Evidence map›Paper›PMID 42538544›Full record

ArticleInternational journal of emergency medicine2026

Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study.

Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand, Ali OmraniNava, Masoud Shahabian, Somayeh Rajabzadeh, Seyyed Mohsen Azizi

Abstract read
In one paragraph

Article in International journal of emergency medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Mehdi Arzani ShamsabadiDepartment of Emergency Medicine, Faculty of Medicine, AJA University of Medical Sciences, Tehran, Iran.
Roya VatankhahMedical Sciences Education Research Center, Mashhad University of Medical Sciences, Mashhad, Iran.
Hasan JalilvandDepartment of Emergency Medicine, Faculty of Medicine, AJA University of Medical Sciences, Tehran, Iran.
Ali OmraniNavaDepartment of Emergency Medicine, Faculty of Medicine, AJA University of Medical Sciences, Tehran, Iran.
Masoud ShahabianClinical Research Development Unit, Besat Hospital, AJA University of Medical Sciences, Tehran, Iran.
Somayeh RajabzadehDepartment of E-Learning in Medical Sciences, Smart University of Medical Sciences, Tehran, Iran. somaye.rajabzade@yahoo.com.
Seyyed Mohsen AziziMedical Education and Development Center, Arak University of Medical Sciences, Arak, Iran.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundTimely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes.

methodsThis descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)-ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1-using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson's Chi-square test, McNemar's test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28).

resultsA total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000).

conclusionLarge language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.

Indexed as

Clinical vignettesDefinitive diagnosesDifferential diagnosesEmergency medicine specialistsLarge language models

Identifiers

PMID42538544
PMCPMC13425980

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.