ArticlePloS one2023
Increasing transparency in machine learning through bootstrap simulation and shapely additive explanations.
Article in PloS one, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 76 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
76 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Prevalence of metabolic syndrome in patients with inflammatory bowel disease: a meta-analysis on a global scale.Journal of health, population, and nutrition · 2025Pooled it
- Artificial Intelligence in Malnutrition: A Systematic Literature Review.Advances in nutrition (Bethesda, Md.) · 2024Pooled it
- Revolutionizing Sleep Medicine: The Impact of Machine Learning on Diagnosis, Treatment, and Personalized Care.Health science reports · 2026Article
- Explainable ensemble machine learning for predicting diabetes mellitus and identifying key risk factors: a population-based study in northern Bangladesh.Scientific reports · 2026Article
- An interpretable machine learning model predicts frailty risk in middle-aged and older adults with gastrointestinal disease: a longitudinal study.Scientific reports · 2026Article
- AI-derived constrained conditional model for screening marker genes through integrated high-throughput transcriptome big data.BMC medical research methodology · 2026Article
- Development of a web platform for predicting fall risk in cardiovascular patients using machine learning.Scientific reports · 2026Article
- Explainable machine learning for depression risk prediction in adults with obesity: development of an online tool.BMC medical informatics and decision making · 2026Article
- Development and Validation of an Explainable Machine Learning Model for Identification of Dysphagia in Patients with COPD.International journal of chronic obstructive pulmonary disease · 2026Article
- Identification of key metabolic indicators associated with the comorbidity of ischemic stroke and diabetes mellitus using an optimal interpretable clinlabomics model.Frontiers in cardiovascular medicine · 2026Article
- A Narrative Review of the Correlation Between Comorbidity and Acute Exacerbation of COPD Patients.International journal of chronic obstructive pulmonary disease · 2026Review
- Construction and internal-external validation of a machine learning-based risk prediction model for multidrug resistance in ICU patients with acute exacerbation of chronic obstructive pulmonary disease.Frontiers in medicine · 2026Article
- Radiomics-machine learning model for predicting invasiveness of subcentimeter subsolid lung adenocarcinoma: a validation study with external cohort and SHAP interpretability.Frontiers in oncology · 2026Article
- Interpretable machine-learning prediction of severe myelosuppression in colorectal cancer patients receiving chemotherapy using XGBoost and SHAP: a retrospective study with a web-based calculator.Frontiers in oncology · 2026Article
- Development, validation, and visualization of a machine learning-based predictive model for depression risk in sleep disorder patients.BMC psychiatry · 2025Article
- Machine Learning Models for Cancer Research: A Narrative Review of Bulk RNA-Seq Applications.International journal of molecular sciences · 2025Review
- Sex-specific machine learning models for cardiovascular disease risk prediction in adults aged ≥ 80 years: insights from the Chinese longitudinal healthy longevity survey.BMC geriatrics · 2025Article
- Development of an explainable machine learning asthma prediction model using serum brominated flame retardants in a national population.Clinical and experimental medicine · 2025Article
- Interpretable machine learning model for low bone density screening in older adults using demographic and anthropometric data: findings from 2005 to 2020 NHANES.BMC medical informatics and decision making · 2025Article
- Predictive modelling in times of public health emergencies: patients' non-transport decisions during the COVID-19 pandemic.BMC emergency medicine · 2025Article
16 more citing papers are in PubMed but not listed here.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
Abstract
Machine learning methods are widely used within the medical field. However, the reliability and efficacy of these models is difficult to assess, making it difficult for researchers to identify which machine-learning model to apply to their dataset. We assessed whether variance calculations of model metrics (e.g., AUROC, Sensitivity, Specificity) through bootstrap simulation and SHapely Additive exPlanations (SHAP) could increase model transparency and improve model selection. Data from the England National Health Services Heart Disease Prediction Cohort was used. After comparison of model metrics for XGBoost, Random Forest, Artificial Neural Network, and Adaptive Boosting, XGBoost was used as the machine-learning model of choice in this study. Boost-strap simulation (N = 10,000) was used to empirically derive the distribution of model metrics and covariate Gain statistics. SHapely Additive exPlanations (SHAP) to provide explanations to machine-learning output and simulation to evaluate the variance of model accuracy metrics. For the XGBoost modeling method, we observed (through 10,000 completed simulations) that the AUROC ranged from 0.771 to 0.947, a difference of 0.176, the balanced accuracy ranged from 0.688 to 0.894, a 0.205 difference, the sensitivity ranged from 0.632 to 0.939, a 0.307 difference, and the specificity ranged from 0.595 to 0.944, a 0.394 difference. Among 10,000 simulations completed, we observed that the gain for Angina ranged from 0.225 to 0.456, a difference of 0.231, for Cholesterol ranged from 0.148 to 0.326, a difference of 0.178, for maximum heart rate (MaxHR) ranged from 0.081 to 0.200, a range of 0.119, and for Age ranged from 0.059 to 0.157, difference of 0.098. Use of simulations to empirically evaluate the variability of model metrics and explanatory algorithms to observe if covariates match the literature are necessary for increased transparency, reliability, and utility of machine learning methods. These variance statistics, combined with model accuracy statistics can help researchers identify the best model for a given dataset.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.