ArticleBMC public health2026
Occupational and socioeconomic predictors of myocardial infarction and coronary heart disease: a machine learning analysis.
Article in BMC public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
Abstract
backgroundCardiovascular disease (CVD) remains a leading cause of morbidity and mortality worldwide. Although traditional cardiovascular risk models primarily rely on biomedical factors, socioeconomic and occupational characteristics are increasingly recognized as important correlates of cardiovascular health. However, applying machine learning to population-based survey data raises methodological concerns, particularly reverse causation and post-diagnosis information leakage.
methodsWe conducted a cross-sectional analysis using data from the 2023 Behavioral Risk Factor Surveillance System (BRFSS). The analytic objective was classification of prevalent myocardial infarction (MI) or coronary heart disease (CHD) rather than prospective risk prediction. Four supervised machine learning algorithms (logistic regression, decision tree, random forest, and gradient boosting) were evaluated. To address potential label leakage, we implemented two model variants: an inclusive model incorporating all available predictors (Model A) and a model excluding post-diagnosis proxy variables such as medication use, functional limitations, and disability indicators (Model B). Model performance was assessed using receiver operating characteristic area under the curve (ROC-AUC), precision-recall area under the curve (PR-AUC), precision, recall, and F1 score.
resultsGradient boosting demonstrated the strongest discriminative performance among the evaluated models. In the inclusive setting, the model achieved a ROC-AUC of 0.867 and a PR-AUC of 0.389, with a best F1 score of 0.433. After removal of post-diagnosis variables, performance remained robust (ROC-AUC = 0.858; PR-AUC = 0.372; F1 = 0.418), suggesting that predictive capacity was not driven solely by downstream disease indicators. Feature importance analyses showed that socioeconomic and employment-related variables remained prominent predictors in Model B alongside established clinical risk factors.
conclusionsMachine learning models can effectively classify prevalent MI and CHD using large-scale survey data even after explicit mitigation of post-diagnosis information leakage. Socioeconomic and occupational characteristics appear to function primarily as contextual correlates rather than causal determinants of CVD. These findings highlight the value of interpretable machine learning approaches for the population-level classification of prevalent cardiovascular disease while underscoring the limitations inherent to cross-sectional data.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.