Evidence map›Paper›PMID 42755637›Full record

ArticleFrontiers in oncology2026

Performance evaluation of large language models in bladder cancer patient education Q&A: a cross-sectional study.

Dian Wan, Youwen Li, Zheng Dong, Chen Dai, Jinghe Ye, Song Li, Sunlu Jiang

Abstract read
In one paragraph

Article in Frontiers in oncology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Dian Wan *Department of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Youwen Li *Department of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Zheng DongDepartment of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Chen DaiCollege of Clinical Medicine, Hubei University of Science and Technology, Xianning, China.
Jinghe YeDepartment of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Song LiDepartment of Urology, Huanggang Central Hospital, Affiliated Huanggang Hospital, Hubei University of Science and Technology, Huanggang, China.
Sunlu JiangDepartment of Minimally Invasive Intervention, Hubei Cancer Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Bladder cancer ranks among the most prevalent urological tumors worldwide, with its global incidence continuing to rise steadily. Although patient education materials (PEMs) play a crucial role in enhancing disease comprehension and supporting joint clinical decision-making, current online resources frequently surpass the readability thresholds recommended for the general public. Large language models (LLMs) hold promise for health communication, yet no systematic assessment has been conducted regarding their feasibility and trustworthiness specifically for bladder cancer patient education. Objective: This study aimed to systematically benchmark five leading LLMs in producing question-and-answer content for bladder cancer science popularization, with a particular focus on readability, informational quality, and appropriateness for patient education. Methods: In this cross-sectional simulation study, 20 common patient questions covering five disease domains were compiled. On January 15, 2026, each question was submitted identically to five publicly available LLMs (Doubao, DeepSeek, Kimi, Gemini, and ChatGPT). Readability was evaluated using seven conventional metrics. Two independent pharmacists, blinded to model identity, rated the responses using the Chinese version of the Patient Education Materials Assessment Tool for print materials (C-PEMAT-P) and the Global Quality Score (GQS). Additionally, two independent clinical specialists assessed factual accuracy and alignment with the Chinese Bladder Cancer Diagnosis and Treatment Guidelines (2024 edition) employing a 4-point scale. Cohen's kappa was used to determine inter-rater reliability. Results: ChatGPT, DeepSeek, and Doubao outperformed Kimi and Gemini on both C-PEMAT and GQS (all P < 0.001), indicating superior understandability, actionability, and overall quality. Across all models, median C-PEMAT scores ranged from 8 to 10, suggesting broadly acceptable suitability for patient education. Readability varied significantly by content domain, with treatment-oriented texts showing the highest complexity. ChatGPT achieved the best alignment with clinical guidelines. No model produced harmful advice or directly contradicted guideline recommendations. Traditional readability measures correlated weakly with GQS, whereas C-PEMAT showed a moderate positive correlation (r = 0.34). Conclusion: Current mainstream LLMs demonstrate initial potential for generating educational content on bladder cancer, albeit with considerable heterogeneity across models. Disease-specific evaluation instruments for patient education materials are more effective than general readability formulas in reflecting perceived quality. Our results advocate for a prudent, assistive role of LLMs in health communication under a human-AI collaborative model.

Indexed as

artificial intelligencebladder cancerlarge language modelspatient educationreadability

Identifiers

PMID42755637
PMCPMC13581702

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.