Evidence map›Paper›PMID 39875808›Full record

ArticleBMC medical research methodology2025

Artificial intelligence methods applied to longitudinal data from electronic health records for prediction of cancer: a scoping review.

Victoria Moglia, Owen Johnson, Gordon Cook, Marc de Kamps, Lesley Smith

Abstract readScoping Review
In one paragraph

Article in BMC medical research methodology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. Review
  11. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Victoria MogliaSchool of Computing, University of Leeds, Woodhouse Lane, Leeds, LS2 9JT, UK. scvcm@leeds.ac.uk.
Owen JohnsonSchool of Computing, University of Leeds, Woodhouse Lane, Leeds, LS2 9JT, UK.
Gordon CookLeeds Institute of Clinical Trials Research, University of Leeds, Clarendon Way, Leeds, LS2 9NL, UK.
Marc de KampsSchool of Computing, University of Leeds, Woodhouse Lane, Leeds, LS2 9JT, UK.
Lesley SmithLeeds Institute of Clinical Trials Research, University of Leeds, Clarendon Way, Leeds, LS2 9NL, UK.

Funding

UK Research and Innovation EP/S024336/1/
6 · The paper itself

Abstract

backgroundEarly detection and diagnosis of cancer are vital to improving outcomes for patients. Artificial intelligence (AI) models have shown promise in the early detection and diagnosis of cancer, but there is limited evidence on methods that fully exploit the longitudinal data stored within electronic health records (EHRs). This review aims to summarise methods currently utilised for prediction of cancer from longitudinal data and provides recommendations on how such models should be developed.

methodsThe review was conducted following PRISMA-ScR guidance. Six databases (MEDLINE, EMBASE, Web of Science, IEEE Xplore, PubMed and SCOPUS) were searched for relevant records published before 2/2/2024. Search terms related to the concepts "artificial intelligence", "prediction", "health records", "longitudinal", and "cancer". Data were extracted relating to several areas of the articles: (1) publication details, (2) study characteristics, (3) input data, (4) model characteristics, (4) reproducibility, and (5) quality assessment using the PROBAST tool. Models were evaluated against a framework for terminology relating to reporting of cancer detection and risk prediction models.

resultsOf 653 records screened, 33 were included in the review; 10 predicted risk of cancer, 18 performed either cancer detection or early detection, 4 predicted recurrence, and 1 predicted metastasis. The most common cancers predicted in the studies were colorectal (n = 9) and pancreatic cancer (n = 9). 16 studies used feature engineering to represent temporal data, with the most common features representing trends. 18 used deep learning models which take a direct sequential input, most commonly recurrent neural networks, but also including convolutional neural networks and transformers. Prediction windows and lead times varied greatly between studies, even for models predicting the same cancer. High risk of bias was found in 90% of the studies. This risk was often introduced due to inappropriate study design (n = 26) and sample size (n = 26).

conclusionThis review highlights the breadth of approaches to cancer prediction from longitudinal data. We identify areas where reporting of methods could be improved, particularly regarding where in a patients' trajectory the model is applied. The review shows opportunities for further work, including comparison of these approaches and their applications in other cancers.

Indexed as

Artificial IntelligenceEarly Detection of CancerElectronic Health RecordsNeoplasmsHumansLongitudinal StudiesReproducibility of ResultsArtificial intelligenceCancerHealth dataLongitudinal dataMachine learningTemporalTime-series

Identifiers

PMID39875808
PMCPMC11773903

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.