Evidence map›Paper›PMID 42106717›Full record

ArticleBMC public health2026

Occupational and socioeconomic predictors of myocardial infarction and coronary heart disease: a machine learning analysis.

Jihyun Noh, Yonghwan Kim, Chaeseong Lim, Su-Yeon Cho, Yujin Kwon, Eun-Soo Lee, Won Kyu Kim, Yun Hak Kim, Kihun Kim

Abstract read
In one paragraph

Article in BMC public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Jihyun Noh *School of Medicine, Pusan National University, Yangsan, 50612, Republic of Korea.
Yonghwan Kim *College of Medicine, Inje University, Busan, 47392, Republic of Korea.
Chaeseong LimDepartment of Internal Medicine, Gachon University Gil Medical Center, Incheon, 21565, Republic of Korea.
Su-Yeon ChoNatural Product Research Center, Korea Institute of Science and Technology, Gangneung, 25451, Republic of Korea.
Yujin KwonDepartment of Convergence Medicine, Yonsei University Wonju College of Medicine, Wonju, 26426, Republic of Korea.
Eun-Soo LeeDepartment of Occupational and Environmental Medicine, Pusan National University Yangsan Hospital, Yangsan, 50612, Republic of Korea.
Won Kyu Kim *Natural Product Research Center, Korea Institute of Science and Technology, Gangneung, 25451, Republic of Korea. wkkim@kist.re.kr.
Yun Hak Kim *Research Institute for Convergence of Biomedical Science and Technology, Pusan National University Yangsan Hospital, Yangsan, 50612, Republic of Korea. yunhak10510@pusan.ac.kr.
Kihun Kim *Department of Occupational and Environmental Medicine, Pusan National University Yangsan Hospital, Yangsan, 50612, Republic of Korea. kihun7603@pnuyh.co.kr.

Funding

Korea Health Industry Development Institute RS-2025-02223691Korea Institute for Advancement of Technology RS-2025-02214034Ministry of Food and Drug Safety RS-2024-00332024National Research Foundation of Korea RS-2023-00223764National Research Foundation of Korea RS-2024-00453708
6 · The paper itself

Abstract

backgroundCardiovascular disease (CVD) remains a leading cause of morbidity and mortality worldwide. Although traditional cardiovascular risk models primarily rely on biomedical factors, socioeconomic and occupational characteristics are increasingly recognized as important correlates of cardiovascular health. However, applying machine learning to population-based survey data raises methodological concerns, particularly reverse causation and post-diagnosis information leakage.

methodsWe conducted a cross-sectional analysis using data from the 2023 Behavioral Risk Factor Surveillance System (BRFSS). The analytic objective was classification of prevalent myocardial infarction (MI) or coronary heart disease (CHD) rather than prospective risk prediction. Four supervised machine learning algorithms (logistic regression, decision tree, random forest, and gradient boosting) were evaluated. To address potential label leakage, we implemented two model variants: an inclusive model incorporating all available predictors (Model A) and a model excluding post-diagnosis proxy variables such as medication use, functional limitations, and disability indicators (Model B). Model performance was assessed using receiver operating characteristic area under the curve (ROC-AUC), precision-recall area under the curve (PR-AUC), precision, recall, and F1 score.

resultsGradient boosting demonstrated the strongest discriminative performance among the evaluated models. In the inclusive setting, the model achieved a ROC-AUC of 0.867 and a PR-AUC of 0.389, with a best F1 score of 0.433. After removal of post-diagnosis variables, performance remained robust (ROC-AUC = 0.858; PR-AUC = 0.372; F1 = 0.418), suggesting that predictive capacity was not driven solely by downstream disease indicators. Feature importance analyses showed that socioeconomic and employment-related variables remained prominent predictors in Model B alongside established clinical risk factors.

conclusionsMachine learning models can effectively classify prevalent MI and CHD using large-scale survey data even after explicit mitigation of post-diagnosis information leakage. Socioeconomic and occupational characteristics appear to function primarily as contextual correlates rather than causal determinants of CVD. These findings highlight the value of interpretable machine learning approaches for the population-level classification of prevalent cardiovascular disease while underscoring the limitations inherent to cross-sectional data.

Indexed as

Coronary DiseaseMachine LearningMyocardial InfarctionOccupationsAgedBehavioral Risk Factor Surveillance SystemBoosting Machine Learning AlgorithmsClassification AlgorithmsCross-Sectional StudiesFemaleHumansMaleMiddle AgedPrediction AlgorithmsPredictive Learning ModelsRandom ForestCardiovascular diseaseMachine learningOccupational factorsPublic healthSocial determinants of health

Identifiers

PMID42106717
PMCPMC13326371

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.