Evidence map›Paper›PMID 42260516›Full record

ArticleBMC musculoskeletal disorders2026

Comparison of three large language models in postoperative rehabilitation question answering after anterior cruciate ligament reconstruction based on expert ratings.

Tiange Xia, Yijie Zhang, Shaoshuo Li, Yi Zhou, Wenyu Tian, Zhiwei Jiang, Yang Shao, Jianwei Wang

Abstract readComparative Study
In one paragraph

Article in BMC musculoskeletal disorders, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Tiange Xia *Nanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China.
Yijie Zhang *Nanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China.
Shaoshuo Li *Wuxi Hospital of Traditional Chinese Medicine, No. 8 Zhongnan West Road, Wuxi, 214071, China.
Yi ZhouNanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China.
Wenyu TianNanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China.
Zhiwei JiangNanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China.
Yang ShaoWuxi Hospital of Traditional Chinese Medicine, No. 8 Zhongnan West Road, Wuxi, 214071, China. wxzy074@njucm.edu.cn.
Jianwei WangNanjing University of Chinese Medicine, No. 138 Xianlin Avenue, Nanjing, 210023, China. wxzy006@njucm.edu.cn.

Funding

Jiangsu Province Traditional Chinese Medicine Science and Technology Development Program QN202322National Natural Science Foundation of China 82274546National Natural Science Foundation of China 82405520
6 · The paper itself

Abstract

backgroundLarge language models (LLMs) are increasingly used by patients to obtain health information. Postoperative rehabilitation after anterior cruciate ligament reconstruction (ACLR) has distinct phase boundaries and safety considerations. Therefore, responses should be not only clear and understandable, but also medically accurate, safe, and stage-fit. This study compared the performance of three publicly accessible LLMs in standardized post-ACLR rehabilitation question answering.

methodsThis was a standardized, blinded, expert-rated comparative evaluation study. On a single prespecified data collection day in March 2026, 30 English-language rehabilitation questions were submitted separately to GPT-5.4, Doubao, and MiniMax-M2.7. The questions covered five postoperative rehabilitation phases. Responses were anonymized and randomly reordered before blinded rating by five orthopaedic clinicians across five domains: Accuracy, Safety, Stage-fit, Completeness, and Understandability. Paired non-parametric tests, effect size analyses, intraclass correlation coefficients, and linear mixed-effects modelling were used for statistical analysis.

resultsA total of 90 model-generated responses and 450 expert rating records were included. Overall scores differed significantly among the three models (Friedman χ² = 46.067, P < 0.001; Kendall's W = 0.768). GPT-5.4 achieved the highest overall score (4.61 ± 0.13), followed by MiniMax-M2.7 (4.53 ± 0.19), whereas Doubao had the lowest score (3.86 ± 0.29). GPT-5.4 performed best in Accuracy, Safety, and Stage-fit; MiniMax-M2.7 achieved the highest score for Completeness; and Doubao achieved the highest mean score for Understandability. Inter-rater agreement was good [ICC(3,k) = 0.893], and sensitivity analysis supported the primary findings.

conclusionsThe three models showed distinct rating profiles in standardized single-turn post-ACLR rehabilitation question answering. Evaluation of patient-facing rehabilitation information should not rely solely on linguistic fluency, but should prioritize medical accuracy, safety, and Stage-fit. These findings provide preliminary benchmark evidence in a phase-sensitive rehabilitation setting, but they should not be interpreted as evidence supporting clinical implementation, clinician substitution, or patient benefit.

Indexed as

Anterior Cruciate Ligament InjuriesAnterior Cruciate Ligament ReconstructionLarge Language ModelsHumansSurveys and QuestionnairesAnterior cruciate ligament reconstructionExpert evaluationLarge language modelsPatient educationPostoperative rehabilitation

Identifiers

PMID42260516
PMCPMC13471468

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.