Evidence map›Paper›PMID 42364231›Full record

ArticleActa orthopaedica et traumatologica turcica2026

Gonarthrosis Advisor vs ChatGPT-5: quality and readability of artificial intelligence-generated patient education for knee osteoarthritis.

Süleyman Kaan Öner, Nihat Demirhan Demirkiran, Enes Alptekin Canlı, Arda Bilir

Abstract readComparative Study
In one paragraph

Article in Acta orthopaedica et traumatologica turcica, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Süleyman Kaan ÖnerDepartment of Orthopaedics and Traumatology, Kütahya Health Sciences University Kütahya City Hospital, Kütahya, Türkiye.
Nihat Demirhan DemirkiranDepartment of Orthopaedics and Traumatology, Kütahya Health Sciences University Kütahya City Hospital, Kütahya, Türkiye.
Enes Alptekin CanlıDepartment of Orthopaedics and Traumatology, Kütahya Health Sciences University Kütahya City Hospital, Kütahya, Türkiye.
Arda BilirDepartment of Orthopaedics and Traumatology, Kütahya Health Sciences University Kütahya City Hospital, Kütahya, Türkiye.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

objectiveThe primary objective of this study is to compare the quality and readability of patient education materials generated by a general-purpose large language model (ChatGPT-5) versus a guideline-based, fine-tuned model (the Gonarthrosis Advisor). The study aims to quantify the performance gains achieved through domain-specific, fine-tuning, and reinforcement learning using osteoarthritis clinical guidelines.  Methods: Thirty frequently asked patient questions regarding knee osteoarthritis were compiled from Google's "People Also Ask" feature and outpatient clinical observations in May 2025. Responses were generated in Turkish by both the Gonarthrosis Advisor and ChatGPT-5 to reflect real-world patient education materials. Content quality was assessed by 2 independent orthopedic surgeons who were blinded to model identity to minimize bias. Both reviewers were co-authors of the article yet did not participate in model construction or data analysis. The assessments utilized the DISCERN instrument, a validated 16-item measure for evaluating the reliability and quality of treatment-related information. Readability was analyzed using the Flesch-Kincaid Grade Level (FKGL), Flesch Reading Ease Score (FRES), and Turkish specific indices (Ateşman and Çakır-Demir). For comparability, English-based indices were applied to translated versions of the responses, whereas Turkish indices were applied to the original texts. All responses were anonymized and randomized prior to evaluation. Model identifiers were removed, and each response was presented in a standardized format to ensure blinding of reviewers. Inter-rater reliability was measured using Cronbach's α. Normality assumptions were tested, and Wilcoxon signed-rank tests were used for statistical comparisons. As no human subjects or personal data were involved, ethical approval was not required.  Results: Mean DISCERN scores corresponded to the "good" category (66.4) for the Gonarthrosis Advisor and the 'moderate' category (54.2) for ChatGPT-5, according to established cut-off thresholds. The Gonarthrosis Advisor achieved significantly higher DISCERN scores than ChatGPT-5 (66.4 ± 4.8 vs. 54.2 ± 5.6; P < .001) with high inter-rater reliability (Cronbach's α = 0.86). Readability metrics favored the Gonarthrosis Advisor across all indices: lower FKGL (7.8 ± 0.7 vs. 9.6 ± 0.9) and higher FRES (54.3 ± 3.4 vs. 46.7 ± 3.7), Ateşman (92.0 ± 4.2 vs. 84.3 ± 4.9), and Çakır-Demir (111.7 ± 5.1 vs. 106.9 ± 5.4) scores (all P <.001).  Conclusion: Fine-tuning large language models with guideline-based content and reinforcement learning improves the quality, neutrality, and accessibility of artificial intelligence-generated patient education materials, offering a scalable tool to enhance health literacy and support shared decision-making in knee osteoarthritis care.    Cite this article as: Öner SK, Demirkiran ND, Canlı EA, Bilir A. Gonarthrosis Advisor vs ChatGPT-5: quality and readability of AI-generated patient education for knee osteoarthritis. Acta Orthop Traumatol Turc., 2026, 60(2), 0618, doi: 10.5152/j.aott.2026.25618.

Indexed as

Artificial IntelligenceComprehensionOsteoarthritis, KneePatient Education as TopicGenerative Artificial IntelligenceHumansLarge Language ModelsReproducibility of ResultsTurkey

Identifiers

PMID42364231
PMCPMC13295167

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.