ArticleCureus2026
Evaluating the Accuracy, Readability, and Consistency of Artificial Intelligence Models in Patient Education for Rheumatoid Arthritis, Osteoarthritis, and Psoriatic Arthritis.
Article in Cureus, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
INTRODUCTION AND
aimArtificial intelligence (AI) chatbots are increasingly used by patients to obtain medical information before seeking clinical care; however, the accuracy, readability, and consistency of AI-generated information in rheumatology remain uncertain. This study aimed to assess the readability, accuracy, and consistency of large language models (LLMs) generated patient education content for rheumatoid arthritis (RA), osteoarthritis (OA), and psoriatic arthritis (PsA).
methodsFrom August 18, 2025, to August 24, 2025, three standardized patient-facing questions per disease were submitted daily to three LLMs, specifically ChatGPT (San Francisco, CA: OpenAI), Google Gemini (Mountain View, CA: Google LLC), and OpenEvidence (Cambridge, MA: OpenEvidence Inc.), with histories cleared between submissions. Readability (Hemingway grade level), word count, accuracy (on a 1-5 scale), and day-to-day consistency (Jaccard similarity) were measured. Responses from each model's most consistent day were accuracy-rated by rheumatology fellows and attendings.
resultsA total of 189 responses were collected. No model consistently met the American Medical Association (AMA)/National Institutes of Health (NIH)-recommended reading levels for sixth through eighth grade. OpenEvidence produced the most technical content (≥17th grade, indicating postbaccalaureate readability), while ChatGPT and Gemini averaged 11.9-12.0 grade. Gemini generated the longest responses (>500 words). ChatGPT showed the highest day-to-day stability (range: <0.07), Gemini moderate variability, and OpenEvidence both the widest range (0.24) and the highest average similarity (0.383). Accuracy ratings varied as follows: OpenEvidence generally scored higher for RA and PsA, while ChatGPT and Gemini were similar across diseases. OA responses showed minimal difference. Best and worst responses mirrored these trends.
conclusionsCurrent LLMs generate rheumatology information above the recommended reading levels. ChatGPT was the most consistent, Gemini the most detailed, and OpenEvidence the most technical. Persistent barriers to readability highlight the need for health-literacy-optimized AI communication.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.