Evidence map›Paper›PMID 36821544›Full record

ArticlePloS one2023

Increasing transparency in machine learning through bootstrap simulation and shapely additive explanations.

Alexander A Huang, Samuel Y Huang

Abstract read
In one paragraph

Article in PloS one, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 76 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
76citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

76 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. A Narrative Review of the Correlation Between Comorbidity and Acute Exacerbation of COPD Patients.International journal of chronic obstructive pulmonary disease · 2026
    Review
  12. Article
  13. Article
  14. Article
  15. Article
  16. Review
  17. Article
  18. Article
  19. Article
  20. Article

16 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Alexander A HuangDepartment of Statistics and Data Science, Cornell University, Ithaca, New York, United States of America.ORCID 0000-0003-4970-4968
Samuel Y HuangDepartment of Statistics and Data Science, Cornell University, Ithaca, New York, United States of America.ORCID 0000-0003-3663-004X

Funding

The Northwestern Summer Research Program for Medical StudentsT35DK126628 · NIDDK · NORTHWESTERN UNIVERSITY AT CHICAGO · PI Daniela P Ladner · 2021 to 2026
$302k
NIDDK NIH HHS T35 DK126628
6 · The paper itself

Abstract

Machine learning methods are widely used within the medical field. However, the reliability and efficacy of these models is difficult to assess, making it difficult for researchers to identify which machine-learning model to apply to their dataset. We assessed whether variance calculations of model metrics (e.g., AUROC, Sensitivity, Specificity) through bootstrap simulation and SHapely Additive exPlanations (SHAP) could increase model transparency and improve model selection. Data from the England National Health Services Heart Disease Prediction Cohort was used. After comparison of model metrics for XGBoost, Random Forest, Artificial Neural Network, and Adaptive Boosting, XGBoost was used as the machine-learning model of choice in this study. Boost-strap simulation (N = 10,000) was used to empirically derive the distribution of model metrics and covariate Gain statistics. SHapely Additive exPlanations (SHAP) to provide explanations to machine-learning output and simulation to evaluate the variance of model accuracy metrics. For the XGBoost modeling method, we observed (through 10,000 completed simulations) that the AUROC ranged from 0.771 to 0.947, a difference of 0.176, the balanced accuracy ranged from 0.688 to 0.894, a 0.205 difference, the sensitivity ranged from 0.632 to 0.939, a 0.307 difference, and the specificity ranged from 0.595 to 0.944, a 0.394 difference. Among 10,000 simulations completed, we observed that the gain for Angina ranged from 0.225 to 0.456, a difference of 0.231, for Cholesterol ranged from 0.148 to 0.326, a difference of 0.178, for maximum heart rate (MaxHR) ranged from 0.081 to 0.200, a range of 0.119, and for Age ranged from 0.059 to 0.157, difference of 0.098. Use of simulations to empirically evaluate the variability of model metrics and explanatory algorithms to observe if covariates match the literature are necessary for increased transparency, reliability, and utility of machine learning methods. These variance statistics, combined with model accuracy statistics can help researchers identify the best model for a given dataset.

Indexed as

AlgorithmsNeural Networks, ComputerComputer SimulationHumansMachine LearningReproducibility of Results

Identifiers

PMID36821544
PMCPMC9949629

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.