Evidence map›Paper›PMID 42700049›Full record

ArticleMedicine2026

Comparative evaluation of large language models for text-based diagnostic reasoning in fetal central nervous system MRI: A retrospective single-center diagnostic accuracy study.

Runze Yu, Miao Peng, Huiying Li, Simeng Liu, Tong Su, Hanjie Guan, Yufeng Shen, Yuhui Deng, Deli Zhao

Abstract readComparative Study
In one paragraph

Article in Medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Runze YuThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Miao PengThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Huiying LiHarbin Red Cross Central Hospital, Harbin, Heilongjiang Province, China.
Simeng LiuThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Tong SuThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Hanjie GuanThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Yufeng ShenThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Yuhui DengThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.
Deli ZhaoThe Sixth Affiliated Hospital of Harbin Medical University, Harbin, Heilongjiang Province, China.ORCID 0000-0001-6193-6361

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Fetal magnetic resonance imaging (MRI) plays an important role in evaluating prenatal central nervous system (CNS) abnormalities, but expert interpretation and counseling remain challenging. Large language models (LLMs) may support text-based diagnostic reasoning, yet their performance in fetal CNS MRI has not been well characterized. The study aimed to compare ChatGPT, Gemini, and DeepSeek in text-based diagnostic reasoning for fetal CNS MRI cases. This retrospective single-center study included 85 fetal MRI cases with postnatal diagnostic confirmation. Each case was converted into a standardized text summary containing MRI findings and essential clinical information, without image input. The 3 LLMs generated ranked differential diagnoses, diagnostic reasoning, management recommendations, prognostic counseling points, and clinical cautions. Anonymized outputs were independently evaluated under blinded conditions for principal diagnosis matching, diagnostic accuracy, clinical reasoning quality, clinical suggestion utility, and communication style and safety. DeepSeek achieved the highest principal diagnosis matching rate (66/85, 77.6%), followed by Gemini (59/85, 69.4%) and ChatGPT (52/85, 61.2%). The overall difference was significant (Cochran's Q = 6.2553, P = .0438), although no pairwise comparison remained significant after Holm correction. DeepSeek showed the highest diagnostic accuracy and reasoning quality, whereas Gemini performed best in clinical suggestion utility and communication style and safety. Exploratory correlation analyses suggested model-specific associations among evaluation metrics. LLM performance in fetal CNS MRI text-based reasoning was multidimensional and model-dependent. These findings support further supervised evaluation of LLMs as text-based decision-support tools but do not support autonomous clinical diagnosis or counseling.

Indexed as

Central Nervous SystemLarge Language ModelsMagnetic Resonance ImagingPrenatal DiagnosisFemaleGenerative Artificial IntelligenceHumansPregnancyRetrospective Studiesartificial intelligenceclinical decision supportfetal diseaseslarge language modelsmagnetic resonance imaging

Identifiers

PMID42700049
PMCPMC13549554

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.