Evidence map›Paper›PMID 42412236›Full record

ArticleArchives of orthopaedic and trauma surgery2026

Artificial intelligence advancements for orthopaedic clinical reasoning: longitudinal assessment of newer models (ChatGPT-5, Grok-3, Gemini 2.5 Flash) compared to clinicians.

Suzen Agharia, Shayan Soroush, Daniel Ameen, Yushy Zhou

Abstract readComparative Study
In one paragraph

Article in Archives of orthopaedic and trauma surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Suzen AghariaFaculty of Medicine, Nursing and Health Sciences, Monash University, Clayton, Australia. suzen.agharia@monash.edu.
Shayan SoroushFaculty of Medicine, Nursing and Health Sciences, Monash University, Clayton, Australia.
Daniel AmeenFaculty of Medicine, Nursing and Health Sciences, Monash University, Clayton, Australia.
Yushy ZhouDepartment of Orthopaedic Surgery, St. Vincent's Hospital, Melbourne, Australia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionThis descriptive study aimed to longitudinally evaluate the performance of contemporary large language models - ChatGPT-5, Gemini 2.5 Flash, and Grok-3 - on orthopaedic clinical multiple-choice tasks, benchmarked against pooled clinician consensus. A secondary aim was to assess whether recent advances in generative AI translated into improved alignment with clinician consensus compared with previous AI models. MATERIALS AND

methodsA total of 97 multiple-choice clinical cases spanning eight orthopaedic subspecialties were sourced from OrthoBullets and previously benchmarked against aggregated responses from thousands of practising clinicians. Using identical methodology to our 2023 study of ChatGPT-3.5, ChatGPT-4, and Bard, each model was prompted with standardised case stems and response options. The primary outcome was the proportion of AI responses matching the most popular clinician response; secondary analyses assessed agreement within 10% and 20% of clinician consensus, performance on 'controversial' (< 25% margin) questions, and inter-model concordance using Cohen's kappa coefficients.

resultsGemini 2.5 Flash achieved the highest alignment with clinician consensus (69.1%), followed by Grok-3 (66.0%) and ChatGPT-5 (58.8%). None of the LLMs refused to respond to any prompts, representing a reduction from 7.2% from our 2023 study. Subspecialty analysis demonstrated that Gemini 2.5 Flash performed best in Hand and Paediatric domains, while Grok-3 excelled in Reconstruction, Trauma, and 'controversial' cases. Inter-model agreement was highest between Grok-3 and Gemini 2.5 Flash (κ = 0.678), indicating improved consistency compared with prior-generation systems.

conclusionsContemporary LLMs can be promising adjuncts for orthopaedic education by simulating peer reasoning and offering structured explanations in non-critical settings. Despite incremental gains in reasoning capability compared to previous AI models, contemporary LLMs remain unsuitable for independent clinical use. Future research should develop hybrid clinician-AI workflows and longitudinal benchmarks to distinguish true reasoning improvements from memorisation.

Indexed as

Artificial IntelligenceClinical ReasoningOrthopedicsGenerative Artificial IntelligenceHumansLarge Language ModelsLongitudinal StudiesBenchmarkingClinician consensusGenerative AILarge language modelsOrthopaedic decision-making

Identifiers

PMID42412236
PMCPMC13342123

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.