Evidence map›Paper›PMID 42688174›Full record

ArticleFrontiers in physiology2026

Patient education for neuromyelitis optica spectrum disorder using large language models: combining expert assessment and real-world patient interaction.

Chen Li, Yuting Hu, Xiaoyan Wang, Ruoyi Yao, Jiahao Ye, Shanshan Diao, Yanhui Xiao, Yingming Wang, Mingying Lai, Weihua Yang

Abstract read
In one paragraph

Article in Frontiers in physiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Chen Li *Department of Ophthalmology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Yuting Hu *Department of Ophthalmology and Optometry, Fujian Medical University, Fuzhou, China.
Xiaoyan Wang *Department of Ophthalmology and Optometry, Fujian Medical University, Fuzhou, China.
Ruoyi YaoDepartment of Ophthalmology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Jiahao YeDepartment of Ophthalmology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Shanshan DiaoDepartment of Neurology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Yanhui XiaoDepartment of Ophthalmology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Yingming WangDepartment of Ophthalmology, The First Affiliated Hospital of Soochow University, Suzhou, China.
Mingying LaiShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, China.
Weihua YangShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objectives: This study aimed to compare the educational performance of six mainstream LLMs for neuromyelitis optica spectrum disorder (NMOSD) and evaluated patient satisfaction during real-world interactions. Methods: This study was conducted from March to April 2026. In the first Phase, Twenty NMOSD-related questions derived from clinical guidelines and patient concerns were submitted to six LLMs (ChatGPT-5.4, Gemini-3.1-pro, Claude-4.6-Sonnet, DeepSeek-3.2, Kimi-2.5, and Qwen-3.5-plus). Responses were anonymized and independently evaluated by three neuro-ophthalmology specialists using Likert framework assessing accuracy, completeness, readability, safety, and humanity. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC). In the second Phase, the three best-performing models were subsequently evaluated through real-world interactions with ten NMOSD patients, and satisfaction scores were analyzed using linear mixed-effects models. Results: A total of 120 chatbot responses were evaluated. With a comprehensive evaluation, significant differences were observed across all assessment domains. Gemini-3.1-pro achieved the highest scores for accuracy and safety, while Qwen-3.5-plus demonstrated superior completeness and humanity. DeepSeek-3.2 generated the most accessible responses, exhibiting the lowest reading difficulty score. However, its completeness advantage should be interpreted with caution, as it may be partially influenced by its longer response length. Inter-rater reliability was good, with single-measure ICC values ranging from 0.535 to 0.759, and average-measure ICC values ranging from 0.775 to 0.904. In patient interactions, Qwen-3.5-plus achieved the highest satisfaction score, significantly outperforming Gemini-3.1-pro and DeepSeek-3.2. Although all LLMs demonstrate superior performance in patient education, the real-world interaction with patient needs to pay attention. Conclusions: LLMs demonstrate considerable potential for NMOSD patient education but exhibit variability across educational dimensions, and require further validation in larger cohorts. These findings highlight the importance of selecting LLMs according to specific patient education goals and underscore the importance of clinicians in rare disease counseling.

Indexed as

artificial intelligencecentral nervous systemhealth educationlarge language modelNMOSD

Identifiers

PMID42688174
PMCPMC13534020

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.