ArticleBMC musculoskeletal disorders2026
Comparison of three large language models in postoperative rehabilitation question answering after anterior cruciate ligament reconstruction based on expert ratings.
Article in BMC musculoskeletal disorders, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
Abstract
backgroundLarge language models (LLMs) are increasingly used by patients to obtain health information. Postoperative rehabilitation after anterior cruciate ligament reconstruction (ACLR) has distinct phase boundaries and safety considerations. Therefore, responses should be not only clear and understandable, but also medically accurate, safe, and stage-fit. This study compared the performance of three publicly accessible LLMs in standardized post-ACLR rehabilitation question answering.
methodsThis was a standardized, blinded, expert-rated comparative evaluation study. On a single prespecified data collection day in March 2026, 30 English-language rehabilitation questions were submitted separately to GPT-5.4, Doubao, and MiniMax-M2.7. The questions covered five postoperative rehabilitation phases. Responses were anonymized and randomly reordered before blinded rating by five orthopaedic clinicians across five domains: Accuracy, Safety, Stage-fit, Completeness, and Understandability. Paired non-parametric tests, effect size analyses, intraclass correlation coefficients, and linear mixed-effects modelling were used for statistical analysis.
resultsA total of 90 model-generated responses and 450 expert rating records were included. Overall scores differed significantly among the three models (Friedman χ² = 46.067, P < 0.001; Kendall's W = 0.768). GPT-5.4 achieved the highest overall score (4.61 ± 0.13), followed by MiniMax-M2.7 (4.53 ± 0.19), whereas Doubao had the lowest score (3.86 ± 0.29). GPT-5.4 performed best in Accuracy, Safety, and Stage-fit; MiniMax-M2.7 achieved the highest score for Completeness; and Doubao achieved the highest mean score for Understandability. Inter-rater agreement was good [ICC(3,k) = 0.893], and sensitivity analysis supported the primary findings.
conclusionsThe three models showed distinct rating profiles in standardized single-turn post-ACLR rehabilitation question answering. Evaluation of patient-facing rehabilitation information should not rely solely on linguistic fluency, but should prioritize medical accuracy, safety, and Stage-fit. These findings provide preliminary benchmark evidence in a phase-sensitive rehabilitation setting, but they should not be interpreted as evidence supporting clinical implementation, clinician substitution, or patient benefit.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.