Evidence map›Paper›PMID 42829621›Full record

ArticleJournal of experimental orthopaedics2026

Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study.

Anita Széll, Yinan Yu, Jacob F Oeding, Felix C Oettl, Adam Piasecki, Keti Dalla, Tobias Siöland, Mathias Hård Af Segerstad, Fredrik Olsen, Peter Larsson and 3 more

Abstract read
In one paragraph

Article in Journal of experimental orthopaedics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Anita SzéllDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.ORCID https://orcid.org/0009-0007-7202-4860
Yinan YuDepartment of Computer Science and Engineering Chalmers University of Technology Gothenburg Sweden.
Jacob F OedingDepartment of Orthopaedics, Institute of Clinical Science Sahlgrenska Academy Gothenburg Sweden.
Felix C OettlDepartment of Orthopaedics, Institute of Clinical Science Sahlgrenska Academy Gothenburg Sweden.
Adam PiaseckiDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Keti DallaDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Tobias SiölandDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Mathias Hård Af SegerstadDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Fredrik OlsenDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Peter LarssonDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.
Stefano ZaffagniniClinica Ortopedica e Traumatologica II IRCCS Istituto Ortopedico Rizzoli Bologna Italy.
Kristian SamuelssonDepartment of Orthopaedics, Institute of Clinical Science Sahlgrenska Academy Gothenburg Sweden.
Fredrik HessulfDepartment of Anesthesiology and Intensive Care Medicine, Institute of Clinical Sciences Sahlgrenska Academy Gothenburg Sweden.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Purpose: Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study. Methods: In this controlled prompting study, 34 orthopaedic anaesthesia questions spanning seven clinical subdomains were independently answered by three LLMs (GPT-3.5-turbo, GPT-4o and GPT-5.2) and three expert anesthesiologists. GPT-4o and GPT-5.2 each generated responses with and without access to relevant clinical practice guidelines. All responses were anonymized and evaluated in blinded pairwise comparisons by an independent guideline-informed LLM-as-a-judge (GPT-5.2) using the same guideline material as the reference standard. The judge recorded preference, response consistency and confidence. Intra-model consistency was assessed from repeated independent generations of each question. Results: All LLMs were preferred over human experts in pairwise comparisons without guideline augmentation, with win rates of 67.6% (GPT-3.5-turbo), 79.4% (GPT-4o) and 95.1% (GPT-5.2). Performance was improved by guideline augmentation, most notably for GPT-4o (91.2%, +11.8 percentage points), while GPT-5.2 approached ceiling performance (97.1%). Response consistency varied between models. GPT-4o showed the highest baseline consistency (80.4%), whereas GPT-5.2 demonstrated no fully contradictory outputs but showed greater partial variability. Guideline augmentation numerically improved GPT-5.2 consistency (66.7% to 80.4%) and reduced inter-model differences. Directed qualitative analysis suggested that reviewer preference for later GPT generations was associated with greater completeness, explicit clinical reasoning and guideline-oriented responses. Conclusion: Successive LLM generations demonstrated progressively improved performance in orthopaedic anaesthesia reasoning. Guideline-augmented prompting enhanced response quality, particularly for intermediate models. Guideline-informed LLM-as-a-judge evaluation appears promising for comparative assessment but requires further validation. Level of Evidence: NA.

Indexed as

GPTlarge language models (LLMs)orthopaedic anaesthesiaprompt engineering

Identifiers

PMID42829621
PMCPMC13633554

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.