ArticleClinical rheumatology2026
ChatGPT-4 vs. DeepSeek-V3: a comparative study of response quality, reliability, usefulness, and readability for exercise and rehabilitation strategies in patients with ankylosing spondylitis.
Article in Clinical rheumatology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introductionThe study assesses the quality, readability, reliability, and usefulness of exercise-related information generated by two large language models (LLMs), ChatGPT-4 and DeepSeek-V3, in response to frequently asked questions by patients with ankylosing spondylitis (AS).
methodThis cross-sectional comparative study developed a structured assessment framework using a set of exercise and rehabilitation-related questions, distributed across four key domains: exercise and physical activity (C1; 33 items), posture and mobility (C2; 6 items), breathing and pulmonary health (C3; 6 items), and general topics (C4; 5 items). Information quality was assessed using the modified DISCERN (mDISCERN) tool, while content reliability was evaluated with the Reliability Score and perceived usefulness was measured using the Usefulness Score. Readability was assessed using the Flesch Reading Ease (FRE) scale. Three independent physiotherapists with expertise in rheumatologic rehabilitation independently evaluated the responses.
resultsIn total score comparisons, DeepSeek-V3 achieved significantly higher scores than ChatGPT-4 on the mDISCERN (4(3-4) vs. 3(3-3); p < 0.001), reliability (5(5-6) vs. 5(4-5); p < 0.001), and usefulness (6(5-6) vs. 5(5-6); p < 0.001). Domain-specific analysis showed higher usefulness scores for DeepSeek-V3 in C1 (p = 0.004), C2 (p = 0.019), and C4 (p = 0.005). Mean FRE scores were 30.4 ± 14.37 for ChatGPT-4 and 28.77 ± 17.77 for DeepSeek-V3, both classified as very difficult (p > 0.05).
conclusionThis study highlighted that responses generated by DeepSeek-V3 related to AS were generally more accurate and demonstrated greater reliability compared to those produced by ChatGPT-4. However, the complex language used by both LLMs may reduce accessibility for patients with limited health literacy. These limitations highlight the importance of healthcare professional oversight in exercise planning. Key Points • DeepSeek-V3 provided more accurate and reliable responses than ChatGPT-4 regarding exercise in AS. • Domain-specific analysis showed DeepSeek-V3 was particularly more useful in exercise, posture, and general topics. • Both LLMs generated content with very difficult readability, requiring college-level comprehension. • Healthcare professional supervision is essential when using LLMs in patient education.
Indexed as
Identifiers
41207975What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.