Evidence map›Paper›PMID 41428044›Full record

ArticleEuropean radiology2026

Predicting molecular types of adult-type diffuse gliomas based on MRI reports with large language models.

Pae Sun Suh, Dahyoun Lee, Chang-Bae Bang, Kyunghwa Han, Kyu Sung Choi, Minjae Kim, Ji Eun Park, Na-Young Shin, Sung Soo Ahn, Seung Hong Choi and 9 more

Abstract readMulticenter Study
PubMed Publisher
In one paragraph

Article in European radiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

19 authors.

Pae Sun Suh *Department of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea.
Dahyoun Lee *Department of Biomedical Systems Informatics, Yonsei University College of Medicine, Seoul, Korea.
Chang-Bae BangDepartment of Psychiatry, Yonsei University College of Medicine, Seoul, Korea.
Kyunghwa HanDepartment of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea.
Kyu Sung ChoiDepartment of Radiology, Seoul National University Hospital, Seoul, Korea.
Minjae KimDepartment of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Seoul, Korea.
Ji Eun ParkDepartment of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Seoul, Korea.
Na-Young ShinDepartment of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea.
Sung Soo AhnDepartment of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea.
Seung Hong ChoiDepartment of Radiology, Seoul National University Hospital, Seoul, Korea.
Ho Sung KimDepartment of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Seoul, Korea.
Seung-Koo LeeDepartment of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea.
Jong Hee ChangDepartment of Neurosurgery, Yonsei University College of Medicine, Seoul, Korea.
Se Hoon KimDepartment of Pathology, Yonsei University College of Medicine, Seoul, Korea.
Martha Foltyn-DumitruDivision for Computational Radiology & Clinical AI (CCIBonn.ai), Clinic for Neuroradiology, University Hospital Bonn, Bonn, Germany.
Seng Chan YouDepartment of Biomedical Systems Informatics, Yonsei University College of Medicine, Seoul, Korea.
Philipp VollmuthDivision for Computational Radiology & Clinical AI (CCIBonn.ai), Clinic for Neuroradiology, University Hospital Bonn, Bonn, Germany.
Byung-Hoon KimDepartment of Biomedical Systems Informatics, Yonsei University College of Medicine, Seoul, Korea. egyptdj@yuhs.ac.
Yae Won ParkDepartment of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Korea. yaewonpark@yuhs.ac.ORCID http://orcid.org/0000-0001-8907-5401

Funding

Ministry of Education NRF-2022R1I1A1A01069589Ministry of Science and ICT, South Korea RS-2024-00509289Yonsei University College of Medicine 6-2023-0072
6 · The paper itself

Abstract

objectivesTo evaluate the performance of large language models (LLMs) in predicting molecular types of adult-type diffuse gliomas according to the 2021 WHO classification using MRI radiology reports. MATERIALS AND

methodsThis retrospective study included 2169 patients diagnosed with adult-type diffuse gliomas (294 oligodendrogliomas, 295 IDH-mutant astrocytomas, and 1580 IDH-wildtype glioblastomas) between July 2005 and March 2024 from four hospitals in Asia and Europe. Seven proprietary and open-source LLMs were assessed: GPT-4o-mini, GPT-4.1-mini, Llama 3.1 8B, Llama 3.1 70B, Qwen2.5 7B, Deepseek-r1 8B, and Mistal 7B. The performance of LLMs in classifying molecular types was compared based on the provision of relevant knowledge of glioma imaging findings (knowledge-based vs. naïve prompt). The impact of radiologists' subspecialization in neuro-oncology, report quality, and reporting language on LLMs' performance was also evaluated.

resultsLLMs achieved significantly higher (naïve vs. knowledge-based; GPT-4o-mini, 77.0% vs. 79.1%, p < 0.001; Qwen2.5 7B, 75.9% vs. 79.5%, p < 0.001; Deepseek-r1 8B, 66.0% vs. 73.2%, p < 0.001) or comparable accuracy (GPT-4.1-mini, 78.7% vs. 78.6%; Llama 3.1 70B, 78.0% vs. 78.1%; Mistral 7B, 58.4% vs. 57.4%) using knowledge-based prompt compared to naïve prompt, except for Llama 3.1 8B (65.4% vs. 44.6%, p < 0.001). Differences in accuracy were more pronounced in smaller-sized LLMs. Additionally, the accuracy was significantly higher with reports by neuro-oncology specialists and high-quality reports in all LLMs (p < 0.001).

conclusionsLLMs may provide preoperative information on the tumor types of adult-type diffuse gliomas from MRI reports by providing relevant knowledge in the prompt. Informative and descriptive reports could further enhance LLMs' performance. KEY POINTS: Question Our study aimed to evaluate large language models' (LLMs) ability to efficiently predict molecular types of adult-type diffuse gliomas according to the 2021 WHO classification. Findings Larger models generally showed better accuracy and were less sensitive to domain-specific knowledge. Their performance improved when using high-quality, longer reports or reports by neuro-oncology specialists. Clinical relevance These findings highlight the potential role of LLMs in predicting glioma molecular types, underscoring the importance of informative and descriptive reports in enhancing their performance.

Indexed as

Brain NeoplasmsGliomaMagnetic Resonance ImagingAdultAgedFemaleHumansLanguageLarge Language ModelsMaleMiddle AgedRetrospective StudiesArtificial intelligenceGliomaLarge language modelMagnetic resonance imaging

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.