Evidence map›Paper›PMID 42428067›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Large language models for cancer registry abstraction: a real-world evaluation across models, variables, and cancer types.

Joshua T Fuchs, Matthew J Satusky, Peter J Leese, Subhadeep Nag, Isaiah W Zipple, Chris D Baggett, Sydney Lash, Katherine Reeder-Hayes, William A Wood, Cara T Johnson and 6 more

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

16 authors.

Joshua T FuchsNorth Carolina Translational and Clinical Sciences Institute, University of North Carolina at Chapel Hill, Chapel Hill, USA.ORCID 0000-0003-2775-3304
Matthew J SatuskyRenaissance Computing Institute, University of North Carolina, Chapel Hill, NC.
Peter J LeeseNorth Carolina Translational and Clinical Sciences Institute, University of North Carolina at Chapel Hill, Chapel Hill, USA.ORCID 0000-0001-9730-1675
Subhadeep NagNorth Carolina Translational and Clinical Sciences Institute, University of North Carolina at Chapel Hill, Chapel Hill, USA.
Isaiah W ZippleLineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Chris D BaggettLineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.ORCID 0000-0002-9563-203X
Sydney LashCarolina Health Informatics Program, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Katherine Reeder-HayesDivision of Oncology, Department of Medicine, UNC School of Medicine, Chapel Hill, North Carolina.
William A WoodDepartment of Medicine, University of North Carolina at Chapel Hill School of Medicine, Chapel Hill.
Cara T JohnsonNorth Carolina Translational and Clinical Sciences Institute, University of North Carolina at Chapel Hill, Chapel Hill, USA.
Claire CritchleyDepartment of Epidemiology, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Ashok K KrishnamurthyRenaissance Computing Institute, University of North Carolina, Chapel Hill, NC.
Jennifer Elston LafataLineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Caroline A ThompsonDepartment of Epidemiology, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Melissa A TroesterDepartment of Epidemiology, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Emily R PfaffDepartment of Medicine, University of North Carolina at Chapel Hill School of Medicine, Chapel Hill.

Funding

North Carolina Translational and Clinical Sciences Institute (NC TraCS)UM1TR004406 · NCATS · UNIV OF NORTH CAROLINA CHAPEL HILL · PI NICHOLAS J SHAHEEN · 2023 to 2026
$37.5M
NCATS NIH HHS UM1 TR004406
6 · The paper itself

Abstract

Cancer registries enable cancer surveillance at the population level. These registries require significant human-time to read through many different parts of the electronic health record, including structured data and lengthy, free-text clinical reports, to abstract values for hundreds of required variables. Large language models (LLMs) offer the possibility to significantly improve this process by supporting and speeding up cancer registry data abstraction. However, it is unclear how well these models perform at real-world cancer registry abstraction involving multiple cancer types and large patient volumes. Here, we evaluate five foundational LLMs for their ability to reliably abstract cancer registry variables. We leverage hospital cancer registry data from a large regional health system as the ground truth and use LLMs to abstract from clinical reports eight registry variables for 5,939 patients with seven different cancer types. We use a zero-shot prompting strategy to compare LLM ability on commonly abstracted cancer variables with different data types. The results show that larger and more advanced models (Claude Sonnet 4.5, GPT-OSS-120b, GPT-OSS-20b) generally outperform smaller models (Gemma 12b, LLaMA 3.1 8b). The best performing models show F1 scores around 0.8 for cancer registry variables with low cardinality (grade, summary stage, laterality), with only slightly lower F1 scores for variables with high cardinality (primary site, regional nodes examined, regional nodes positive). On the more complex task of precise date extraction, all models showed decreased performance on both diagnosis and treatment dates (exact accuracy ~0.55 for the best performing models), which increased to ~0.85 for a tolerance within ±30 days. These results quantify the performance of various models as well as the potential and limitations of LLMs in cancer registry abstraction tasks.

Identifiers

PMID42428067
PMCPMC13345411

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.