Evidence map›Paper›PMID 42376461›Full record

ArticleFrontiers in cell and developmental biology2026

Risk-centered benchmarking of large language models for AI-enabled counseling in chronic autoimmune thyroid eye disease.

Fangqin Fei, Lu Xie, Jing Rao, Ziqi Liang, Juan Yang, Yunyun Zou, Ligang Jiang

Abstract read
In one paragraph

Article in Frontiers in cell and developmental biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Fangqin Fei *Department of Endocrinology, The First People's Hospital of Huzhou, Huzhou Normal University, Huzhou, Zhejiang, China.
Lu Xie *Shenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Jing Rao *Shenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Ziqi LiangShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Juan YangShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Yunyun ZouShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Ligang JiangDepartment of Ophthalmology, Quzhou People's Hospital, The Quzhou Affiliated Hospital, Wenzhou Medical University, Wenzhou, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Thyroid eye disease (TED) is a chronic autoimmune inflammatory orbital disease requiring activity assessment, risk stratification, and triage. As patients increasingly consult large language models (LLMs), evidence on their quality and safety for TED counseling remains limited. Methods: We conducted a cross-sectional benchmark using a prespecified 35-question Chinese TED counseling bank covering symptom recognition, activity assessment, treatment, daily management, follow-up, and care-seeking. The bank was developed from guideline-/consensus-derived scenarios, expert discussion, and recurrent patient-inquiry themes, then applied under a unified single-turn protocol. Five web-based LLM chatbots were evaluated: Gemini 3 Pro, ChatGPT-5.2, DeepSeek-V3.1, Doubao, and Qwen3-Max. Systems were accessed through official interfaces in Quzhou, China, during 27-29 December 2025, with identifiers recorded. Automated text analysis extracted output features, and response time was measured. Two blinded expert raters assessed alignment with a guideline-/consensus-informed reference standard using 5-point Likert scales for Accuracy, Logic, Coherence, Safety, and Content Accessibility. Between-model comparisons used repeated-measures methods, and correlations used Spearman analysis. Results: Response time differed significantly across models (Friedman χ Conclusion: LLMs show marked heterogeneity in efficiency, output structure, and clinical quality for TED counseling. Longer or slower responses do not necessarily indicate better performance. The absence of between-model Safety differences should not be interpreted as absolute safety or equivalence. Risk-centered, structured outputs emphasizing red-flag symptoms and care-seeking thresholds warrant validation through multi-turn dialogues, repeated sampling, patient/lay-user evaluation, and finer-grained safety endpoints.

Indexed as

artificial intelligenceautoimmune inflammationchronic ocular diseaselarge language modelsrisk stratificationthyroid eye disease

Identifiers

PMID42376461
PMCPMC13311113

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.