Evidence map›Paper›PMID 40611010›Full record

ArticleBMC pulmonary medicine2025

Construction and validation of a risk prediction model for chronic obstructive pulmonary disease (COPD): a cross-sectional study based on the NHANES database from 2009 to 2018.

Liqin Wang, Shijia Zhang, Zhaohong Gao, Deyou Jiang

Abstract readValidation Study
In one paragraph

Article in BMC pulmonary medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Liqin WangHeilongjiang University of Chinese Medicine, 24 Heping Road, Xiangfang District, Harbin, 150040, Heilongjiang, China.
Shijia ZhangHeilongjiang University of Chinese Medicine, 24 Heping Road, Xiangfang District, Harbin, 150040, Heilongjiang, China.
Zhaohong GaoHeilongjiang University of Chinese Medicine, 24 Heping Road, Xiangfang District, Harbin, 150040, Heilongjiang, China.
Deyou JiangHeilongjiang University of Chinese Medicine, 24 Heping Road, Xiangfang District, Harbin, 150040, Heilongjiang, China. klxz1469@sina.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundChronic obstructive pulmonary disease (COPD) is a major global public health concern, and early screening and identification of high-risk populations are critical for reducing the disease burden. Although several studies have explored the application of machine learning methods in COPD risk prediction, existing models often have limited feature dimensions and insufficient interpretability. Identifying key risk factors and constructing reliable predictive models remain challenges in clinical practice.

objectiveThis study aims to integrate multidimensional features based on data from the National Health and Nutrition Examination Survey (NHANES) and to compare the performance of different machine learning models in COPD risk prediction. The goal is to identify the optimal model and enhance its clinical applicability through interpretability analysis.

methodsThis study utilized data from the NHANES collected between 2009 and 2018. After systematic feature selection and preprocessing, three models were developed: multivariate binary logistic regression, XGBoost, and Multilayer Perceptron (MLP). Model training and evaluation were performed using stratified five-fold cross-validation. Model performance was comprehensively assessed based on accuracy, precision, recall, F1 score, and the area under the receiver operating characteristic curve (AUC). To enhance model transparency, the SHapley Additive Explanations (SHAP) method was employed to interpret key features and their influence trends within the MLP model.

resultsThe MLP model demonstrated the best performance across all evaluation metrics, achieving an average accuracy of 0.937, precision of 0.6624, recall of 0.6535, and F1 score of 0.657 in stratified five-fold cross-validation. The performance gap between the training and testing sets was minimal, indicating no obvious overfitting. SHAP analysis identified smoking years, asthma, age, dietary health status, total protein, red cell distribution width (RDW), BMI, marital status, secondhand smoke exposure, and total bilirubin as important predictive features. Furthermore, dependence plots revealed critical risk inflection points for key continuous variables.

conclusionBased on large-scale and multidimensional feature data, this study constructed a COPD risk prediction model with favorable performance and enhanced interpretability. The findings suggest that the MLP model has the potential to effectively identify individuals at high risk for COPD and may offer value in clinical applications. Future studies are warranted to integrate longitudinal follow-up data and multimodal information to further improve predictive accuracy and clinical interpretability, thereby providing a more robust foundation for early screening and personalized interventions in COPD.

Indexed as

Machine LearningPulmonary Disease, Chronic ObstructiveAdultAgedCross-Sectional StudiesFemaleHumansLogistic ModelsMaleMiddle AgedNutrition SurveysRisk AssessmentRisk FactorsROC CurveUnited StatesCOPDMachine learningNHANES databaseRisk prediction model

Identifiers

PMID40611010
PMCPMC12225070

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.