Evidence map›Paper›PMID 34032855›Full record

ArticleJAMA network open2021

Development and Validation of a Machine Learning Model Using Administrative Health Data to Predict Onset of Type 2 Diabetes.

Mathieu Ravaut, Vinyas Harish, Hamed Sadeghi, Kin Kwan Leung, Maksims Volkovs, Kathy Kornas, Tristan Watson, Tomi Poutanen, Laura C Rosella

Abstract readValidation Study
In one paragraph

Article in JAMA network open, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 49 papers, 4 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
49citing papers in PubMed, 4 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

49 citing papers in PubMed, 4 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Pooled it
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Review
  14. Article
  15. Article
  16. Article
  17. Review
  18. Article
  19. Article
  20. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Mathieu RavautLayer 6 AI, Toronto, Ontario, Canada.
Vinyas HarishDalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.
Hamed SadeghiLayer 6 AI, Toronto, Ontario, Canada.
Kin Kwan LeungLayer 6 AI, Toronto, Ontario, Canada.
Maksims VolkovsLayer 6 AI, Toronto, Ontario, Canada.
Kathy KornasDalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.
Tristan WatsonDalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.
Tomi PoutanenLayer 6 AI, Toronto, Ontario, Canada.
Laura C RosellaDalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: Systems-level barriers to diabetes care could be improved with population health planning tools that accurately discriminate between high- and low-risk groups to guide investments and targeted interventions. Objective: To develop and validate a population-level machine learning model for predicting type 2 diabetes 5 years before diabetes onset using administrative health data. Design, Setting, and Participants: This decision analytical model study used linked administrative health data from the diverse, single-payer health system in Ontario, Canada, between January 1, 2006, and December 31, 2016. A gradient boosting decision tree model was trained on data from 1 657 395 patients, validated on 243 442 patients, and tested on 236 506 patients. Costs associated with each patient were estimated using a validated costing algorithm. Data were analyzed from January 1, 2006, to December 31, 2016. Exposures: A random sample of 2 137 343 residents of Ontario without type 2 diabetes was obtained at study start time. More than 300 features from data sets capturing demographic information, laboratory measurements, drug benefits, health care system interactions, social determinants of health, and ambulatory care and hospitalization records were compiled over 2-year patient medical histories to generate quarterly predictions. Main Outcomes and Measures: Discrimination was assessed using the area under the receiver operating characteristic curve statistic, and calibration was assessed visually using calibration plots. Feature contribution was assessed with Shapley values. Costs were estimated in 2020 US dollars. Results: This study trained a gradient boosting decision tree model on data from 1 657 395 patients (12 900 257 instances; 6 666 662 women [51.7%]). The developed model achieved a test area under the curve of 80.26 (range, 80.21-80.29), demonstrated good calibration, and was robust to sex, immigration status, area-level marginalization with regard to material deprivation and race/ethnicity, and low contact with the health care system. The top 5% of patients predicted as high risk by the model represented 26% of the total annual diabetes cost in Ontario. Conclusions and Relevance: In this decision analytical model study, a machine learning model approach accurately predicted the incidence of diabetes in the population using routinely collected health administrative data. These results suggest that the model could be used to inform decision-making for population health planning and diabetes prevention.

Indexed as

Age of OnsetAlgorithmsDecision Making, Computer-AssistedMachine LearningAdolescentAdultAgedAged, 80 and overChildCohort StudiesDiabetes Mellitus, Type 2Electronic Health RecordsFemaleForecastingHumansIncidence

Identifiers

PMID34032855
PMCPMC8150694

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.