Evidence map›Paper›PMID 41723213›Full record

ArticleScientific reports2026

A comparative analysis of data-driven models for breast cancer survival prediction.

Kasahun Takele, Ding-Geng Chen

Abstract readComparative Study
In one paragraph

Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. A pan-cancer multi-omicFrontiers in artificial intelligence · 2026
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Kasahun TakeleDepartment of Statistics, Haramaya University, Maya, Ethiopia. kastake10@gmail.com.
Ding-Geng ChenDepartment of Statistics, University of Pretoria, Pretoria, South Africa.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Breast cancer is the most frequently diagnosed cancer among women and persists as a societal problem worldwide. It remains a leading cause of cancer associated morbidity and mortality, specifically in low- and middle-income countries where access to timely diagnosis and treatment is often limited. This study aims to compare survival and classical machine learning models for predicting breast cancer survival in Ethiopia to identify approaches that balance predictive accuracy with interpretability. The study utilized retrospective data from 1164 women treated at Tikur Anbesa Specialized Hospital and Hiwot Fana Specialized University Hospital between 2019 and 2024. Methods like Kaplan-Meier estimation, Cox proportional hazards, random survival forests (RSF), DeepSurv, and classical machine learning (SVM, XGBoost, LGBM, and RF) classifiers were used with evaluation metrics such as AUC, C-index, and Integrated Brier Score (IBS). The Shapley additive explanation approach was used to ensure the interpretability of results from models such as RSF, DeepSurv, and random forests (RF). It allowed the identification of important predictors of breast cancer outcome by indicating consistent predictors across models. The findings demonstrated that random survival forest and random forest achieved the highest performance (C-index: 0.754; IBS: 0.091) and (0.729 ± 0.006), respectively, outperforming the other models under consideration. The Shapley Additive Explanations (SHAP) analysis for the RSF model showed that age, tumour size, metastasis, stage, comorbidities, and marital status as the most important predictors of breast cancer survival. Furthermore, the SHAP analysis for the RF model indicated that the higher age category (45 and above), metastasis status (M1), stage four, and larger tumour size contribute a strong influence on predictions. Among the machine learning models, the random forest algorithm effectively identifies the key predictors of breast cancer outcomes. For the survival analysis methods, the RSF offers robust capabilities for handling time-to-event data and censoring, making it well-suited for accurate survival prediction. By combining these approaches, we were able to gain clearer insights and better identify the key factors influencing breast cancer prognosis. This study highlights the value of data-driven methods in helping healthcare professionals identify high-risk patients with greater precision and take timely, informed actions to support their care.

Indexed as

Breast NeoplasmsAdultBoosting Machine Learning AlgorithmsClassification AlgorithmsEthiopiaFemaleHumansKaplan-Meier EstimateMachine LearningMiddle AgedPrediction AlgorithmsPredictive Learning ModelsPrognosisProportional Hazards ModelsRandom ForestRetrospective StudiesBreastCancerC-indexDeepSurvSHAP

Identifiers

PMID41723213
PMCPMC13022381

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.