Evidence map›Paper›PMID 37164379›Full record

SynthesisBMJ (Clinical research ed.)2023

Development and internal-external validation of statistical and machine learning models for breast cancer prognostication: cohort study.

Ash Kieran Clift, David Dodwell, Simon Lord, Stavros Petrou, Michael Brady, Gary S Collins, Julia Hippisley-Cox

Open access · hybridAbstract readMeta-Analysis
In one paragraph

Synthesis in BMJ (Clinical research ed.), 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 69 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
69citing papers in PubMed, 2 pooled it
19.1field-weighted citation impact, top 1% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

69 citing papers in PubMed, 2 syntheses or guidelines pooled it, 84 citations in OpenAlex.

  1. Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. An interpretable breast cancer risk stratification model via multi-omics integration: multi-method development and cross-cohort validation.Clinical & translational oncology : official publication of the Federation of Spanish Oncology Societies and of the National Cancer Institute of Mexico · 2026
    Article
  8. Article
  9. Association between residential greenness and the risk of inflammatory bowel disease in patients with COVID-19: A Korean nationwide cohort study.Saudi journal of gastroenterology : official journal of the Saudi Gastroenterology Association · 2026
    Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Review
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article

9 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors at 2 institutions in 1 country.

Ash Kieran CliftCancer Research UK Oxford Centre, Oxford, UK ashley.clift@phc.ox.ac.uk.ORCID 0000-0002-0061-979X
David DodwellNuffield Department of Population Health, University of Oxford, Oxford, UK.
Simon LordDepartment of Oncology, University of Oxford, Oxford, UK.
Stavros PetrouNuffield Department of Primary Care Health Sciences, Radcliffe Primary Care Building, Radcliffe Observatory Quarter, University of Oxford, Oxford OX2 6GG, UK.
Michael BradyDepartment of Oncology, University of Oxford, Oxford, UK.
Gary S CollinsCentre for Statistics in Medicine, Nuffield Department of Orthopaedics, Rheumatology and Musculoskeletal Sciences, University of Oxford, Oxford, UK.
Julia Hippisley-CoxNuffield Department of Primary Care Health Sciences, Radcliffe Primary Care Building, Radcliffe Observatory Quarter, University of Oxford, Oxford OX2 6GG, UK.
University of Oxford · GBNuffield Orthopaedic Centre · GB

Funding

Cancer Research UK 27294
6 · The paper itself

Abstract

objectiveTo develop a clinically useful model that estimates the 10 year risk of breast cancer related mortality in women (self-reported female sex) with breast cancer of any stage, comparing results from regression and machine learning approaches.

designPopulation based cohort study.

settingQResearch primary care database in England, with individual level linkage to the national cancer registry, Hospital Episodes Statistics, and national mortality registers.

participants141 765 women aged 20 years and older with a diagnosis of invasive breast cancer between 1 January 2000 and 31 December 2020.

main outcome measuresFour model building strategies comprising two regression (Cox proportional hazards and competing risks regression) and two machine learning (XGBoost and an artificial neural network) approaches. Internal-external cross validation was used for model evaluation. Random effects meta-analysis that pooled estimates of discrimination and calibration metrics, calibration plots, and decision curve analysis were used to assess model performance, transportability, and clinical utility.

resultsDuring a median 4.16 years (interquartile range 1.76-8.26) of follow-up, 21 688 breast cancer related deaths and 11 454 deaths from other causes occurred. Restricting to 10 years maximum follow-up from breast cancer diagnosis, 20 367 breast cancer related deaths occurred during a total of 688 564.81 person years. The crude breast cancer mortality rate was 295.79 per 10 000 person years (95% confidence interval 291.75 to 299.88). Predictors varied for each regression model, but both Cox and competing risks models included age at diagnosis, body mass index, smoking status, route to diagnosis, hormone receptor status, cancer stage, and grade of breast cancer. The Cox model's random effects meta-analysis pooled estimate for Harrell's C index was the highest of any model at 0.858 (95% confidence interval 0.853 to 0.864, and 95% prediction interval 0.843 to 0.873). It appeared acceptably calibrated on calibration plots. The competing risks regression model had good discrimination: pooled Harrell's C index 0.849 (0.839 to 0.859, and 0.821 to 0.876, and evidence of systematic miscalibration on summary metrics was lacking. The machine learning models had acceptable discrimination overall (Harrell's C index: XGBoost 0.821 (0.813 to 0.828, and 0.805 to 0.837); neural network 0.847 (0.835 to 0.858, and 0.816 to 0.878)), but had more complex patterns of miscalibration and more variable regional and stage specific performance. Decision curve analysis suggested that the Cox and competing risks regression models tested may have higher clinical utility than the two machine learning approaches.

conclusionIn women with breast cancer of any stage, using the predictors available in this dataset, regression based methods had better and more consistent performance compared with machine learning approaches and may be worthy of further evaluation for potential clinical use, such as for stratified follow-up.

Indexed as

Breast NeoplasmsCohort StudiesEnglandFemaleHumansMachine LearningRisk Assessment

Identifiers

PMID37164379
PMCPMC10170264
OpenAlexW4376118166

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.