ArticleBMJ open2025
How have generic large language models progressed in their ability to write clinic letters and provide accurate management plans in the virtual fracture clinic?
Article in BMJ open, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
1 citing paper in PubMed.
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
objectiveTo explore whether large language models (LLMs), Generative Pre-trained Transformer (GPT)-3, GPT-3.5 and GPT-4 can autonomously manage a virtual fracture clinic (VFC) as a marker of their efficacy in an emergency department and with simple orthopaedic trauma. SETTING AND
participantsSimulated UK VFC workflow.
design11 clinical scenarios were generated, and GPT-4, GPT-3.5 and GPT-3 were prompted to write clinic letters and management plans.
main outcome measuresThe Readable Tool was used to assess the clarity of letters. Six independent orthopaedic surgeons then evaluated the accuracy of letters and management plans.
resultsReadability was compared using the Flesch-Kincaid grade level: GPT-4: 9.11 (SD 0.98); GPT-3.5: 8.77; GPT-3: 8.47, and the Flesch readability ease: GPT-4: 56.3; GPT-3.5: 58.2; GPT-3: 59.3. Surgeon-rated accuracy comparisons indicated that GPT-4 exhibited the highest accuracy for management plans (9.08/10 (95% CI 8.25 to 9.9)). This represents a statistically significant progression in the capacity of a LLM to provide accurate management plans compared with GPT-3 at 6.84 (95% CI 5.41 to 8.27) and GPT-3.5 at 7.63 (95% CI 7.23 to 8.13) (p<0.0001).
conclusionsLLMs can produce high-quality, readable clinical letters for common VFC presentations, and GPT-4 can generate management plans to aid clinicians in their administration. With clinician oversight, appropriately trained LLMs could meaningfully reduce routine administrative work. However, while the results of this study are promising, further evaluation of LLMs is required before they can be deemed safe for managing simple orthopaedic scenarios.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.