ArticleScientific reports2026
A comparative analysis of data-driven models for breast cancer survival prediction.
Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
3 citing papers in PubMed.
- Prognostic impact of extensive nodal involvement in small-sized (T1) breast cancer: a population-based and propensity score-matched analysis of the SEER database.Gland surgery · 2026Article
- Pharmacoeconomic evaluation of first-line tislelizumab for extensive-stage small cell lung cancer using a comparative validation of traditional survival and machine learning models.Frontiers in public health · 2026Article
- A pan-cancer multi-omicFrontiers in artificial intelligence · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Breast cancer is the most frequently diagnosed cancer among women and persists as a societal problem worldwide. It remains a leading cause of cancer associated morbidity and mortality, specifically in low- and middle-income countries where access to timely diagnosis and treatment is often limited. This study aims to compare survival and classical machine learning models for predicting breast cancer survival in Ethiopia to identify approaches that balance predictive accuracy with interpretability. The study utilized retrospective data from 1164 women treated at Tikur Anbesa Specialized Hospital and Hiwot Fana Specialized University Hospital between 2019 and 2024. Methods like Kaplan-Meier estimation, Cox proportional hazards, random survival forests (RSF), DeepSurv, and classical machine learning (SVM, XGBoost, LGBM, and RF) classifiers were used with evaluation metrics such as AUC, C-index, and Integrated Brier Score (IBS). The Shapley additive explanation approach was used to ensure the interpretability of results from models such as RSF, DeepSurv, and random forests (RF). It allowed the identification of important predictors of breast cancer outcome by indicating consistent predictors across models. The findings demonstrated that random survival forest and random forest achieved the highest performance (C-index: 0.754; IBS: 0.091) and (0.729 ± 0.006), respectively, outperforming the other models under consideration. The Shapley Additive Explanations (SHAP) analysis for the RSF model showed that age, tumour size, metastasis, stage, comorbidities, and marital status as the most important predictors of breast cancer survival. Furthermore, the SHAP analysis for the RF model indicated that the higher age category (45 and above), metastasis status (M1), stage four, and larger tumour size contribute a strong influence on predictions. Among the machine learning models, the random forest algorithm effectively identifies the key predictors of breast cancer outcomes. For the survival analysis methods, the RSF offers robust capabilities for handling time-to-event data and censoring, making it well-suited for accurate survival prediction. By combining these approaches, we were able to gain clearer insights and better identify the key factors influencing breast cancer prognosis. This study highlights the value of data-driven methods in helping healthcare professionals identify high-risk patients with greater precision and take timely, informed actions to support their care.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.