ArticleJournal of cancer education : the official journal of the American Association for Cancer Education2026
Evaluating the Performance of Large Language Models for Breast Cancer Patient Education: A Comparative Study.
Article in Journal of cancer education : the official journal of the American Association for Cancer Education, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
12 authors.
Funding
Abstract
Breast Cancer necessitates effective patient education. Large language models (LLMs) facilitate patient health consultation, yet their generated medical content may contain misleading and unsafe information. Systematic evaluations of mainstream LLMs for breast cancer health guidance are currently lacking. This study evaluated six LLMs' (ChatGPT-5.4-thinking, Claude-4.6-sonnet, Gemini-3.1-Pro, DeepSeek-V3.2, Doubao-2.2-thinking, and ERNIE 4.5 Turbo) performance in breast cancer consultation via a structured checklist. A set of 61 standardized questions regarding breast cancer was developed based on Google Trends, clinical guidelines, practical experiences, and expert reviews. Responses from each LLM were independently evaluated by three breast cancer experts focusing on quality, accuracy, comprehensiveness, and safety. Besides, four patients independently evaluated the satisfaction and understandability of their selected three questions of interest. This study utilized Bernard's Global Quality Score (GQS) tool to assess quality. Readability was assessed using the Chinese Resource Platform (CRP). Other indicators were evaluated using self-designed questionnaires. Statistical analyses were performed using RStudio. In expert evaluations, ERNIE 4.5 Turbo had the highest descriptive quality score and was among the top-performing models in safety (Bonferroni-adjusted P < 0.05), while several models performed comparably in comprehensiveness. There was no significant difference in accuracy among the models. ChatGPT-5.4-thinking scored significantly lower in safety, and Doubao-2.2-thinking had significantly lower reading difficulty, required age, and Chinese character count (adjusted P < 0.05). In patient evaluations, ERNIE 4.5 Turbo showed the highest descriptive satisfaction and understandability ratings. Six large language models performed strongly in breast cancer question-answering, with ERNIE 4.5 Turbo ranking highest. However, issues like poor readability and unsafe recommendations remain in answers. Future research should prioritize enhancing patient readability to facilitate AI's application in precision cancer health education.
Indexed as
Identifiers
42228312What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.