Evidence map›Paper›PMID 41729569›Full record

ArticleJMIR cancer2026

Evaluation of GPT-5 for Esophageal Cancer Staging Using Fluorodeoxyglucose Positron Emission Tomography Maximum-Intensity Projection Images: Comparative Pilot Study.

Hiroki Maruyama, Yoshitaka Toyama, Yuya Araki, Kentaro Takanami, Masato Ito, Yumi Nakajima, Kei Takase, Takashi Kamei

Abstract readComparative Study
In one paragraph

Article in JMIR cancer, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Hiroki MaruyamaDepartment of Surgery, Graduate School of Medicine, Tohoku University, Sendai, Japan.ORCID 0009-0008-7845-2281
Yoshitaka ToyamaDepartment of Imaging and Anatomy for Groundbreaking Education Collaborative Research, Graduate School of Medicine, Tohoku University, Sendai, Japan.ORCID 0000-0003-0027-9681
Yuya ArakiSchool of Medicine, Tohoku University, Sendai, Japan.ORCID 0009-0005-7035-6680
Kentaro TakanamiDepartment of Diagnostic Radiology, Tohoku University Hospital, Sendai, Japan.ORCID 0000-0002-0098-7760
Masato ItoDepartment of Diagnostic Radiology, Tohoku University Hospital, Sendai, Japan.ORCID 0009-0008-7979-999X
Yumi NakajimaDepartment of Diagnostic Radiology, Tohoku University Hospital, Sendai, Japan.ORCID 0009-0003-2948-2357
Kei TakaseDepartment of Diagnostic Radiology, Tohoku University Hospital, Sendai, Japan.ORCID 0000-0003-0931-9942
Takashi KameiDepartment of Surgery, Graduate School of Medicine, Tohoku University, Sendai, Japan.ORCID 0000-0003-1282-0463

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAccurate esophageal cancer staging relies on

objectiveWe evaluated the diagnostic accuracy of LLMs for staging esophageal cancer using

methodsThis retrospective study included 120 consecutive adult patients who were diagnosed with esophageal squamous cell carcinoma and underwent

resultsThe average accuracy was 41/120 (34%) to 94/120 (78%) for LLMs and 72/120 (60%) to 102/120 (85%) for physicians, with significantly higher accuracy for physicians (P<.05) in the thoracic LN, abdominal LN, and cN stages. Interrater reliability was slight to fair for LLMs (κ: -0.07 to 0.25) and fair to substantial for physicians (κ: 0.27 to 0.74). Matthews Correlation Coefficient scores were consistently higher for physicians (0.28 to 0.75) than for LLMs (-0.07 to 0.32). Among the LLMs, GPT-5 demonstrated the highest overall accuracy, with newer LLMs showing improved diagnostic accuracy when compared with previous models in identifying abdominal LN metastases and cM staging, though they showed weaker consistency for cN staging. For example, in thoracic LN detection, GPT-5 achieved 76/120 (63%) accuracy, whereas other LLMs achieved 72/120 (60%) or lower accuracy.

conclusionsAlthough current LLMs have not yet reached physician-level accuracy in comprehensive staging, recent models show promise in assisting with specific diagnostic tasks.

Indexed as

Esophageal NeoplasmsEsophageal Squamous Cell CarcinomaFluorodeoxyglucose F18Positron-Emission TomographyPositron Emission Tomography Computed TomographyAdultAgedFemaleHumansLarge Language ModelsLymphatic MetastasisMaleMiddle AgedNeoplasm StagingPilot ProjectsRadiopharmaceuticalsFluorodeoxyglucose F18Radiopharmaceuticals18F FDG-PET imagingesophageal cancer stagingfluorodeoxyglucose positron emission tomographygenerative artificial intelligencelarge language modelsLLMsradiology report automation

Identifiers

PMID41729569
PMCPMC12972682

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.