Evidence map›Paper›PMID 41404998›Full record

ArticleKnee surgery, sports traumatology, arthroscopy : official journal of the ESSKA2026

Reasoning-optimised large language models reach near-expert accuracy on board-style orthopaedic exams: A multi-model comparison on 702 multiple-choice questions.

Pedro Diniz, Takuji Yokoe, Felix C Öttl, Hélder Pereira, Rui Henriques, Kristian Samuelsson

Abstract readComparative Study
In one paragraph

Article in Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Pedro DinizDepartment of Orthopaedic Surgery, Centre Hospitalier Universitaire Brugmann, Brussels, Belgium.ORCID https://orcid.org/0000-0001-9234-7041
Takuji YokoeDivision of Orthopaedic Surgery, Department of Medicine of Sensory and Motor Organs, Faculty of Medicine, University of Miyazaki, Miyazaki, Japan.
Felix C ÖttlDepartment of Orthopaedic Surgery, Balgrist University Hospital, Zürich, Switzerland.
Hélder PereiraOrthopaedic Department, Centro Hospitalar Póvoa de Varzim, Vila do Conde, Portugal.
Rui HenriquesINESC-ID and Instituto Superior Técnico, Universidade de Lisboa, Lisboa, Portugal.
Kristian SamuelssonDepartment of Orthopaedics, Institute of Clinical Sciences, Sahlgrenska Academy, University of Gothenburg, Gothenburg, Sweden.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeThe purpose of this study was to compare the accuracy, calibration, reproducibility and operating cost of seven large language models (LLMs)-including four newer models capable of using advanced reasoning techniques to analyse complex medical information and generate accurate responses-on text-only orthopaedic multiple-choice questions (MCQs) and to quantify gains over GPT-4.

methodsFrom Orthobullets, 702 unique, non-image MCQs (drawn from AAOS Self-Assessment Examinations, Self-Assessment-Based Questions and Orthopaedic In Training Examination-Based Questions banks) were extracted. Each question was submitted to the following LLMs: OpenAI o3, Anthropic Claude Sonnet 4, Claude Opus 4 (with/without 'Extended Thinking') and Google Gemini 2.5 Pro. Additionally, OpenAI's GPT-4, GPT-4o and the open-weight Gemma 3 27B served as comparators. The primary outcome was overall accuracy. The secondary outcomes were topic and difficulty-stratified accuracy, calibration (expected calibration error [ECE] and Brier score), reproducibility (flip rate on a retest question subset), latency, token use and cost. Statistical tests included paired McNemar, Cochran Q, ordinal logistic regression and Fleiss κ (Bonferroni-adjusted α = 0.05).

resultsGPT-4 achieved 69.7% accuracy (95% CI = 66.2-72.9). All four reasoning-optimised models scored ≥14 percentage points higher (p < 3.3 × 10

conclusionsReasoning-optimised LLMs now answer text-based orthopaedic exam questions with high accuracy and substantially better confidence calibration than earlier models. However, persistent stochasticity and large latency-cost disparities may limit clinical deployment. LEVEL OF EVIDENCE: N/A.

Indexed as

Clinical CompetenceEducational MeasurementLanguageOrthopedicsCalibrationHumansLarge Language ModelsReproducibility of Resultsartificial intelligenceclinical decision supportlarge language modelsmedical educationorthopaedic surgery

Identifiers

PMID41404998
PMCPMC12850567

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.