Evidence map›Paper›PMID 42055579›Full record

ArticleThe British journal of general practice : the journal of the Royal College of General Practitioners2026

Using large language models to identify prediagnostic clinical features of ovarian cancer from healthcare records: a population-based case-control study.

Garth Funston, Namu Park, Matthew Thompson, Meliha Yetisgen, Barbara A Goff, Larry Kessler, Fiona M Walter

Abstract read
In one paragraph

Article in The British journal of general practice : the journal of the Royal College of General Practitioners, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. What the codes don't say.The British journal of general practice : the journal of the Royal College of General Practitioners · 2026
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Garth FunstonCentre for Cancer Screening, Prevention and Early Diagnosis, Wolfson Institute of Population Health, Barts and The London School of Medicine and Dentistry, Queen Mary University of London, London, UK g.funston@qmul.ac.uk.
Namu ParkDepartment of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, US.
Matthew ThompsonDepartment of Family Medicine, University of Washington, Seattle, US.
Meliha YetisgenDepartment of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, US.
Barbara A GoffDepartment of Obstetrics and Gynecology, University of Washington, Seattle, WA, US.
Larry KesslerDepartment of Health Systems and Population Health, School of Public Health, University of Washington, Seattle, WA, US.
Fiona M WalterCentre for Cancer Screening, Prevention and Early Diagnosis, Wolfson Institute of Population Health, Barts and The London School of Medicine and Dentistry, Queen Mary University of London, London, UK.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundMost women with ovarian cancer are diagnosed after developing symptoms. However, symptoms are often recorded as free text within electronic health records (EHRs), which is not readily accessible for research.

aimTo use EHRs to examine associations between coded and large language model (LLM)-extracted free-text clinical features with ovarian cancer diagnosis. DESIGN AND

settingPopulation-based case-control study using EHRs and cancer registry data from women attending primary care, outpatient, and emergency clinics associated with the University of Washington, US.

methodIn total, 136 women with ovarian cancer cases (diagnosed 2012-2019) were matched (age, clinic type) to 1360 control participants. Twelve months of prediagnosis coded and free-text data were extracted from EHRs. LLMs were tested on annotated notes, before extracting information on 17 prespecified clinical features. Univariate conditional logistic regression analyses were used to identify clinical features associated with ovarian cancer.

resultsThere were 14 clinical features that were more commonly identified from free text using LLMs than from codes in both the case and control groups. There were 14 features that were significantly associated with ovarian cancer when using codes and LLM-extracted data, but only eight features were significant using codes alone. Using both coded and LLM-extracted data, 11 features had odds ratios >2. Thirteen features were significantly associated when restricting analysis to early-stage (I-II) diagnosis.

conclusionThere is an identifiable ovarian cancer symptom signature within EHRs, with LLM-based natural language processing approaches enabling extraction of key non-coded symptom information. LLMs could support EHR-based research, while LLM-based clinical decision support tools may improve identification of patients with symptoms of possible cancer.

Indexed as

Electronic Health RecordsOvarian NeoplasmsAgedCase-Control StudiesFemaleHumansLarge Language ModelsMiddle AgedRegistriescancerearly diagnosislarge language modelsovarianprimary health care

Identifiers

PMID42055579
PMCPMC13531583

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.