ArticleThe British journal of general practice : the journal of the Royal College of General Practitioners2026
Using large language models to identify prediagnostic clinical features of ovarian cancer from healthcare records: a population-based case-control study.
Article in The British journal of general practice : the journal of the Royal College of General Practitioners, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
- Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses.Journal of imaging · 2026Article
- What the codes don't say.The British journal of general practice : the journal of the Royal College of General Practitioners · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundMost women with ovarian cancer are diagnosed after developing symptoms. However, symptoms are often recorded as free text within electronic health records (EHRs), which is not readily accessible for research.
aimTo use EHRs to examine associations between coded and large language model (LLM)-extracted free-text clinical features with ovarian cancer diagnosis. DESIGN AND
settingPopulation-based case-control study using EHRs and cancer registry data from women attending primary care, outpatient, and emergency clinics associated with the University of Washington, US.
methodIn total, 136 women with ovarian cancer cases (diagnosed 2012-2019) were matched (age, clinic type) to 1360 control participants. Twelve months of prediagnosis coded and free-text data were extracted from EHRs. LLMs were tested on annotated notes, before extracting information on 17 prespecified clinical features. Univariate conditional logistic regression analyses were used to identify clinical features associated with ovarian cancer.
resultsThere were 14 clinical features that were more commonly identified from free text using LLMs than from codes in both the case and control groups. There were 14 features that were significantly associated with ovarian cancer when using codes and LLM-extracted data, but only eight features were significant using codes alone. Using both coded and LLM-extracted data, 11 features had odds ratios >2. Thirteen features were significantly associated when restricting analysis to early-stage (I-II) diagnosis.
conclusionThere is an identifiable ovarian cancer symptom signature within EHRs, with LLM-based natural language processing approaches enabling extraction of key non-coded symptom information. LLMs could support EHR-based research, while LLM-based clinical decision support tools may improve identification of patients with symptoms of possible cancer.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.