ArticleDigital health
An explainable multi-label Diagnostic Prediction Model for
Article in Digital health. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Objectives: Differential diagnosis of pneumonia, tuberculosis, and lung cancer is highly challenging due to overlapping clinical presentations and high comorbidity rates. To overcome the limitations of traditional diagnostic methods, this study developed and validated an explainable machine learning model using proteomic data for multi-label classification of these three diseases. Methods: Bronchoalveolar lavage fluid proteomic data were collected from 358 patients with confirmed lung diseases at the Shandong Provincial Public Health Clinical Center. We constructed a voting ensemble model integrating XGBoost, Random Forest, and Gradient Boosting algorithms based on clinical features, global proteomic statistical features, differentially expressed derived features, and disease-specific biomarker scores. Considering insufficient cancer samples and risk of missed diagnoses, a 1.5-fold weighting strategy was applied to the cancer class to enhance sensitivity. The model was evaluated using 5-fold stratified cross-validation and an independent external cohort of 110 cases. Feature contributions were interpreted using the Shapley Additive exPlanations (SHAP) method and biological significance of key proteins was determined using Gene Ontology functional enrichment analysis. Results: The AUC values were 0.912±0.031 for tuberculosis detection and 0.813±0.037 for cancer detection. Thus, the ensemble model demonstrated excellent performance, significantly outperforming seven baseline models including logistic regression and support vector machines. SHAP analysis identified key protein biomarkers (Cancer: P02775, P61626; Tuberculosis: P0DOX2, P55259). Furthermore, the model achieved an overall accuracy of 86% with the independent external validation cohort. Conclusion: This study established an explainable, multi-label classification model based on proteomics that can provide a valuable reference for the differential diagnosis of complex lung diseases, especially those with comorbidities. The model shows good performance and interpretability, suggesting potential for precise diagnosis of lung diseases.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.