BackPublic Health

Using tree-based ensemble methods to produce a population-based mortality risk score in Ontario, Canada

PLoS OneResearch Authors: Steven Habbous , Peter C. Austin, Shabnam Balamchi, Davood Astaraky, Roozbeh Yousefi, Munaza Chaudhry, Erik HellstenAIIM Authors: Pearl Marks, Amanda ZhongApproved by President Reda RiffiPublication Date: 4/23/2026

Comprehensive Summary

This retrospective, population-based single-center study asked whether tree-based machine learning models could improve prediction of 1-year all-cause mortality in the general adult population compared with traditional statistical methods, using ensemble tree-based machine learning models to perform mortality risk prediction. Researchers analyzed 12,080,801 adult Ontarians alive as of January 1, 2022, using administrative healthcare data, including hospital discharge records, ambulatory care encounters, physician billing claims, laboratory data, cancer registry data, and long-term care assessments collected from provincial databases in Ontario, Canada. Predictors were demographic characteristics, Charlson-style comorbidities, healthcare utilization measures, physician billing indicators, and Activities of Daily Living (ADL) scores. Preprocessing included retaining continuous variables without scaling, assigning missing ADL scores to a separate category, and transforming categorical variables using one-hot encoding. The models tested included logistic regression, random forest, ExtraTrees, AdaBoost, gradient boosting, Newton boosting/XGBoost-style models, CatBoost, and explainable boosting machines, compared primarily against standard logistic regression. The best-performing model was CatBoost, which achieved an AUROC of 0.933, PR-AUC of 0.281, Brier score of 0.0083, and Integrated Calibration Index equivalent (ICIeq) of 0.0003, outperforming logistic regression, which achieved an AUROC of 0.926 and PR-AUC of 0.256. External temporal validation in a 2024 cohort maintained an AUROC of 0.933 with a slightly lower PR-AUC of 0.254 and modest calibration drift. The analysis showed that 1.0% of the cohort died within one year, while 51.3% were female and the mean age was 49.0 years. Diabetes without complications was the most prevalent comorbidity at 4.3%, followed by primary cancer at 2.1% and diabetes with complications at 2.0%. Mortality rates were highest among patients receiving palliative care (36%), residing in long-term care (27%), experiencing dementia (26%), pressure injury (25%), delirium (21%), metastatic cancer (20%), and stage 4–5 chronic kidney disease (16–17%). Secondary analyses included calibration assessment, Kaplan-Meier survival stratification, explainability modeling using feature importance and permutation feature importance, marginal effect estimation, sensitivity analyses evaluating alternate cancer and kidney disease definitions, and comparison of 70/30 train-test splitting with 10-fold cross-validation. Additional results demonstrated that age, outpatient hospital visits, sex, healthcare utilization frequency, ADL score, palliative care status, and hospitalization burden were among the strongest predictors of mortality. Marginal effect analyses showed that palliative care was associated with an absolute mortality risk increase of 4.03% and a relative marginal effect of 438%, while moderate-to-severe liver disease and metastatic cancer increased mortality risk by 2.60% and 1.53%, respectively. Explainable boosting machine analyses further demonstrated that age interacted strongly with multiple clinical variables in predicting mortality risk. Limitations include reliance on administrative coding data, incomplete disease granularity, such as cancer staging and subtype information, limited computational resources preventing exhaustive hyperparameter tuning, and potential algorithmic bias introduced by healthcare utilization variables that may underrepresent underserved populations. Although external validation was performed using a later Ontario cohort, subgroup fairness analyses were limited, and advanced bias mitigation techniques were not implemented. Findings reflect predictive discrimination and calibration within large-scale administrative datasets rather than direct improvements in patient-centered outcomes and therefore do not establish clinical efficacy.

Outcomes and Implications

This study suggests that tree-based ensemble models, particularly CatBoost, can improve population-level mortality prediction and risk adjustment compared with traditional regression-based approaches. In clinical practice, these models could support health system planning, identify high-risk patients for targeted supportive interventions such as palliative care consultation, and improve confounder adjustment in epidemiologic research. At a population level, mortality risk stratification tools could inform resource allocation, surveillance, and proactive care management strategies. However, translation to bedside care requires ongoing recalibration, prospective multicenter validation, and rigorous fairness assessment to ensure equitable performance across sociodemographic groups.

Our mission is to

Connect medicine with AI innovation.

No spam. Only the latest AI breakthroughs, simplified and relevant to your field.