Evidence map›Paper›PMID 42499817›Full record

ArticleDigital health

An explainable multi-label Diagnostic Prediction Model for

Yan Wang, Jie Tan, Xinjun Li, Zhaoyan Yu, Bingzheng Wang, Liang Sun

Abstract read
In one paragraph

Article in Digital health. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Yan WangCollege of Medical Information and Artificial Intelligence, Shandong First Medical University and Shandong Academy of Medical Sciences, Binzhou People's Hospital Affiliated to Shandong First Medical University, Jinan, Shandong, China.
Jie TanDepartment of Respiratory and Critical Care Medicine, The Affiliated Hospital of Inner Mongolia Medical University, Hohhot, China.
Xinjun LiDepartment of Pathology, Binzhou People's Hospital.
Zhaoyan YuFirst Affiliated Hospital, Biomedical Sciences College, Shandong Academy of Medical Sciences, Shandong Medicinal Biotechnology Centre, Shandong First Medical University, Shandong, China.
Bingzheng WangShandong Provincial Key Laboratory of Development and Regeneration, School of Life Science, Shandong University, Shandong, China.
Liang SunCollege of Medical Information and Artificial Intelligence, Shandong First Medical University and Shandong Academy of Medical Sciences, Binzhou People's Hospital Affiliated to Shandong First Medical University, Jinan, Shandong, China.ORCID https://orcid.org/0000-0002-5213-6941

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objectives: Differential diagnosis of pneumonia, tuberculosis, and lung cancer is highly challenging due to overlapping clinical presentations and high comorbidity rates. To overcome the limitations of traditional diagnostic methods, this study developed and validated an explainable machine learning model using proteomic data for multi-label classification of these three diseases. Methods: Bronchoalveolar lavage fluid proteomic data were collected from 358 patients with confirmed lung diseases at the Shandong Provincial Public Health Clinical Center. We constructed a voting ensemble model integrating XGBoost, Random Forest, and Gradient Boosting algorithms based on clinical features, global proteomic statistical features, differentially expressed derived features, and disease-specific biomarker scores. Considering insufficient cancer samples and risk of missed diagnoses, a 1.5-fold weighting strategy was applied to the cancer class to enhance sensitivity. The model was evaluated using 5-fold stratified cross-validation and an independent external cohort of 110 cases. Feature contributions were interpreted using the Shapley Additive exPlanations (SHAP) method and biological significance of key proteins was determined using Gene Ontology functional enrichment analysis. Results: The AUC values were 0.912±0.031 for tuberculosis detection and 0.813±0.037 for cancer detection. Thus, the ensemble model demonstrated excellent performance, significantly outperforming seven baseline models including logistic regression and support vector machines. SHAP analysis identified key protein biomarkers (Cancer: P02775, P61626; Tuberculosis: P0DOX2, P55259). Furthermore, the model achieved an overall accuracy of 86% with the independent external validation cohort. Conclusion: This study established an explainable, multi-label classification model based on proteomics that can provide a valuable reference for the differential diagnosis of complex lung diseases, especially those with comorbidities. The model shows good performance and interpretability, suggesting potential for precise diagnosis of lung diseases.

Indexed as

ensemble learningexplainabilitylung diseasesmachine learningmulti-label classification

Identifiers

PMID42499817
PMCPMC13396575

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.