Evidence map›Paper›PMID 40426209›Full record

ArticleJournal of orthopaedic surgery and research2025

The role of large language models in improving the readability of orthopaedic spine patient educational material.

Melissa Romoff, Madison Brunette, Melanie K Peterson, Sohaib Z Hashmi, Michael S Kim

Abstract read
In one paragraph

Article in Journal of orthopaedic surgery and research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed.

  1. Comparison of responses from google and large language models to the top frequently asked questions on lumbar spinal stenosis: an evaluation of accuracy and completeness.European spine journal : official publication of the European Spine Society, the European Spinal Deformity Society, and the European Section of the Cervical Spine Research Society · 2026
    Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Review
  9. Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Melissa RomoffDepartment of Orthopaedic Surgery, University of California, Irvine, School of Medicine, 101 The City Dr S, Pavilion 3, Building 29 A, Orange, CA, 92868, USA.
Madison BrunetteDepartment of Orthopaedic Surgery, University of California, Irvine, School of Medicine, 101 The City Dr S, Pavilion 3, Building 29 A, Orange, CA, 92868, USA.
Melanie K PetersonDepartment of Orthopaedic Surgery, University of California, Irvine, School of Medicine, 101 The City Dr S, Pavilion 3, Building 29 A, Orange, CA, 92868, USA.
Sohaib Z HashmiDepartment of Orthopaedic Surgery, University of California, Irvine, School of Medicine, 101 The City Dr S, Pavilion 3, Building 29 A, Orange, CA, 92868, USA.
Michael S KimDepartment of Orthopaedic Surgery, University of California, Irvine, School of Medicine, 101 The City Dr S, Pavilion 3, Building 29 A, Orange, CA, 92868, USA. michak14@hs.uci.edu.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionPatient education is crucial for informed decision-making. Current educational materials are often written at a higher grade level than the American Medical Association (AMA)-recommended sixth-grade level. Few studies have assessed the readability of orthopaedic materials such as American Academy of Orthopaedic Surgeons (AAOS) OrthoInfo articles, and no studies have suggested efficient methods to improve readability. This study assessed the readability of OrthoInfo spine articles and investigated the ability of large language models (LLMs) to improve readability.

methodsA cross-sectional study analyzed 19 OrthoInfo articles using validated readability metrics (Flesch-Kincaid Grade Level and Reading Ease). Articles were simplified iteratively in three steps using ChatGPT, Gemini, and CoPilot. LLMs were prompted to summarize text, followed by two clarification prompts simulating patient inquiries. Word count, readability, and accuracy were assessed at each step. Accuracy was rated by two independent reviewers using a three-point scale (3 = fully accurate, 2 = minor inaccuracies, 1 = major inaccuracies). Statistical analysis included one-way and two-way ANOVA, followed by Tukey post-hoc tests for pairwise comparisons.

resultsBaseline readability exceeded AMA recommendations, with a mean Flesch-Kincaid Grade Level of 9.5 and a Reading Ease score of 51.1. LLM summaries provided statistically significant improvement in readability, with the greatest improvements in the first iteration. All three LLMs performed similarly, though ChatGPT achieved statistically significant improvements in Reading Ease scores. Gemini incorporated appropriate disclaimers most consistently. Accuracy remained stable throughout, with no evidence of hallucination or compromise in content quality or medical relevance. DISCUSSION: LLMs effectively simplify orthopaedic educational content by reducing grade levels, enhancing readability, and maintaining acceptable accuracy. Readability improvements were most significant in initial simplification steps, with all models performing consistently. These findings support the integration of LLMs into patient education workflows, offering a scalable strategy to improve health literacy, enhance patient comprehension, and promote more equitable access to medical information across diverse populations.

Indexed as

ComprehensionHealth LiteracyLanguageOrthopedicsPatient Education as TopicSpineTeaching MaterialsCross-Sectional StudiesHumansLarge Language Models

Identifiers

PMID40426209
PMCPMC12117680

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.