ArticlePloS one2026
Using tree-based ensemble methods to produce a population-based mortality risk score in Ontario, Canada.
Article in PloS one, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introductionRisk adjustment is critical in observational epidemiology to control for confounding of the exposure-outcome relationship. Accurate prediction of outcomes, such as mortality, can improve risk adjustment. In the present study, we compared logistic regression with a range of tree-based ensemble methods to predict 1-year mortality in the general population of Ontario, Canada.
methodsOntario adults (age 18 years and older) who were alive as of January 1, 2022 were included. Using a window of up to 3 years, various measures of health and healthcare utilization were captured from administrative databases. To predict 1-year mortality, we applied logistic regression, random forests, extremely randomized trees, adaptive boosting, gradient boosting, extreme gradient boosting, Newton boosting, and CatBoost. All models also included age and sex. Performance was evaluated using the area under the ROC curve (AUROC), the area under the precision-recall curve (PR-AUC), the Brier score, and a quantile-based version of the Integrated Calibration Index (ICI), reported in the 30% test set. Feature importance was assessed using CatBoost's internal model structure, supplemented with permutation feature importance, explainable boosted machines, and marginal effects.
resultsA total of 12,080,801 Ontarians were included and 121,951 (1.0%) died within 1 year. Logistic regression showed excellent discrimination (AUROC 0.926; PR-AUC 0.256) and acceptable calibration (ICI 0.0022). The best model was CatBoost, which had the best discrimination (AUROC 0.933, PR-AUC 0.280) and calibration (ICI 0.0003). In sensitivity analyses of the CatBoost model, including more detailed definitions of cancer (to include its subtype) and chronic kidney disease (defined using serum creatinine instead of diagnostic codes) produced modest improvements in PR-AUC (0.290), along with substantially improved calibration amongst the highest-risk (70-100%) individuals. The most influential model-building feature was age. Residence in long-term care and receipt of palliative care was associated with the largest marginal effects.
conclusionThe machine learning model CatBoost yielded the most accurate predictive model for 1-year mortality using individual comorbidities and additional measures of healthcare utilization for the general population. These findings demonstrate that machine learning methods can enhance risk adjustment efforts in observational studies, leading to more accurate confounder control and better support for health policy and epidemiologic research.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.