Evidence map›Paper›PMID 41586949›Full record

ArticleEuropean radiology experimental2026

Reliability of Gemini 2.5 Pro, ChatGPT 4.1, DeepSeek V3, and Claude Opus 4 in generating standardized CMR protocols.

Răzvan-Andrei Licu, Giuseppe Muscogiuri, Davide Casartelli, Anca Bacârea, Marian Pop, Andra-Maria Licu, Daniele Sferratore, Alessandro Caruso, Marianna Mirchuk, Piotr Tarkowski and 2 more

Abstract read
In one paragraph

Article in European radiology experimental, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Răzvan-Andrei Licu *School of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Giuseppe Muscogiuri *School of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy. g.muscogiuri@gmail.com.ORCID http://orcid.org/0000-0003-4757-2420
Davide CasartelliSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Anca BacâreaDepartment of Pathophysiology, George Emil Palade University of Medicine, Pharmacy, Science and Technology, Târgu Mureș, Romania.
Marian PopDepartment of Radiology, County Emergency Clinical Hospital, Târgu Mureș, Romania.
Andra-Maria LicuDepartment of Radiology, County Emergency Clinical Hospital, Târgu Mureș, Romania.
Daniele SferratoreDepartment of Radiology, ASST Papa Giovanni XXIII, Bergamo, Italy.
Alessandro CarusoSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Marianna MirchukSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Piotr TarkowskiSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Jakub ByczkowskiSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.
Sandro SironiSchool of Medicine and Surgery, University of Milano-Bicocca, Milan, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Artificial intelligence (AI) and large language models (LLMs) are increasingly integrated into radiology, offering new possibilities for advanced imaging techniques, including cardiovascular magnetic resonance (CMR). This proof-of-concept study assessed four high-performing LLMs (Gemini 2.5 Pro, ChatGPT 4.1, DeepSeek V3, and Claude Opus 4) on their ability to generate CMR protocols for 140 hypothetical cardiac cases. AI-generated protocols were compared against a reference standard established by a consensus between two experienced cardiovascular radiologists, following the Society for Cardiovascular Magnetic Resonance (SCMR) recommendations. Descriptive statistics were used to quantify the concordance of LLM-generated sequences with the SCMR guidelines. Statistical agreement was measured using Cohen and Fleiss κ statistics. Gemini 2.5 Pro achieved the highest concordance, aligning with the SCMR guidelines in 71.5% of all evaluated scenarios. Overall, LLMs showed moderate agreement with the SCMR protocols, with Gemini 2.5 Pro again performing best (Cohen κ = 0.55). Agreement was substantial for mandatory CMR sequences (Fleiss κ ∈ [0.69, 0.74]) and predominantly fair for optional sequences. The tested LLMs demonstrate a potential to generate efficient and pathology-adapted CMR protocols. Under expert supervision, this capability could streamline the imaging workflow and help extend CMR to primary healthcare centers through protocol automation. RELEVANCE STATEMENT: The potential of Gemini 2.5 Pro, ChatGPT 4.1, DeepSeek V3, and Claude Opus 4 to suggest pathology-adapted CMR protocols could improve imaging throughput and help to expand access to advanced cardiac diagnostics in primary healthcare centers. KEY POINTS: The tested large language models show potential for generating CMR protocols. Substantial agreement on mandatory CMR sequences promises more efficient examinations. Automation of CMR protocols could help to improve access to this advanced technique outside major medical institutions.

Indexed as

Magnetic Resonance ImagingGenerative Artificial IntelligenceHumansLarge Language ModelsReproducibility of ResultsArtificial intelligenceCardiovascular diseaseLarge language modelsMagnetic resonance imagingRadiology

Identifiers

PMID41586949
PMCPMC12834875

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.