Evidence map›Paper›PMID 41441998›Full record

ArticleEuropean radiology2026

Identification of high-priority radiology reports with unexpected findings using fine-tuned large language models.

Akihiro Umeno, Mizuho Nishio, Hidetoshi Matsuo, Takaaki Matsunaga, Munenobu Nogami, Eisuke Ueshima, Keitaro Sofue, Takamichi Murakami

Abstract read
PubMed Publisher
In one paragraph

Article in European radiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

8 authors.

Akihiro UmenoDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Mizuho NishioDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan. nishiomizuho@gmail.com.ORCID http://orcid.org/0000-0001-5870-0868
Hidetoshi MatsuoDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Takaaki MatsunagaDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Munenobu NogamiDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Eisuke UeshimaDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Keitaro SofueDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.
Takamichi MurakamiDepartment of Radiology, Kobe University Graduate School of Medicine, Chuo-ku, Japan.

Funding

JSPS 22K07665JSPS 23K17229JSPS 23KK0148SIP JPJ01242
6 · The paper itself

Abstract

objectiveThis study aims to evaluate whether large language models (LLMs) can accurately predict the urgency and severity of radiology reports. MATERIALS AND

methodsBased on the recommendations of the Academy of Royal Colleges, we defined radiology reports that include unexpected findings of high urgency or severity as "high-priority (HP) radiology reports." Overall, 1906 radiology reports were used as the training set, and 176 radiology reports were used as the test set, with a balanced ratio of HP to non-HP radiology reports (1:1) in both sets. Four types of LLMs (Llama2 7B, Llama3 8B, Llama3 Elyza 8B, and Llama 3.1 8B) were fine-tuned using four different input settings: (1) findings only, (2) findings + referring department, (3) findings + referring department + clinical diagnosis before examination, and (4) findings + referring department + clinical diagnosis before examination + details of examination request. The fine-tuned LLMs predicted whether each radiology report was HP or not.

resultsAmong the four LLMs, Llama3 Elyza 8B, with inputs comprising findings and the referring department, demonstrated the best performance, achieving PRAUC = 0.962, ROCAUC = 0.968, accuracy = 0.915, sensitivity/recall = 0.932, specificity = 0.898, and F1 = 0.916. Adding a clinical diagnosis before the examination and details of examination requests did not necessarily lead to performance improvement.

conclusionThe fine-tuned LLMs accurately predicted HP radiology reports, suggesting their potential utility in supporting communication regarding radiology reports with high urgency or severity. KEY POINTS: Question This study aims to evaluate whether large language models (LLMs) can accurately predict the high-priority (HP) radiology reports. Findings The fine-tuned best LLM accurately HP radiology reports, achieving PRAUC of 0.962 and ROCAUC of 0.968. Clinical relevance This study demonstrates that fine-tuned LLMs can accurately identify HP radiology reports, potentially improving timely clinical decision-making and enhancing patient safety through faster communication of critical findings.

Indexed as

Radiology Information SystemsHumansLarge Language ModelsDeep learningGenerative AILarge language modelRadiology reportSafety

Identifiers

PMID41441998

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.