ArticleOral health & preventive dentistry2026
Performance of Large Language Models in Oral Cancer Patient Education: An Evaluation of Reliability, Readability, and Patient Communication Quality.
Article in Oral health & preventive dentistry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundAs society is increasingly depending on large language models (LLMs) for health-related questions, it is essential to objectively evaluate the quality and accessibility of the oral cancer information they provide. Although LLMs occupy a growing space in digital health communication, it remains unknown whether the information they generate is both reliable and easy to read.
objectiveThe purpose of this study was to evaluate the reliability and readability, respectively, of the responses generated by four mainstream LLMs (ChatGPT, Gemini, Perplexity, and DeepSeek) to questions related to oral cancer. Specifically, the present authors aimed to evaluate the reliability of responses to common oral cancer questions and assess whether the readability of responses meets established expectations.
methodsTwenty-two commonly asked, patient-orientated oral cancer-related questions were developed through two predefined phases: Google Trends analysis and expert consultation with specialists in oral oncology. Each question was entered as an independent single-turn prompt into four LLMs: ChatGPT-5, Gemini 2.5, Perplexity Pro, and DeepSeek v3.2. The primary outcome was information reliability and quality, assessed using four standardized instruments: the DISCERN questionnaire, the Ensuring Quality Information for Patients (EQIP) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). The secondary outcome was readability, assessed using six established indices: the Automated Readability Index, Flesch Reading Ease Score, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Simple Measure of Gobbledygook.
resultsSignificant differences were observed among the four LLMs in DISCERN, EQIP, and JAMA scores (all P 0.001), whereas no significant difference was found in GQS scores (P = 0.440). Perplexity Pro achieved the highest mean DISCERN score (46.36 ± 4.70), EQIP score (85.00 ± 0.00), GQS score (4.05 ± 0.58), and JAMA score (1.00 ± 0.00). However, all models produced responses above the recommended sixth-grade readability level. The mean FKGL scores ranged from 12.65 ± 3.07 for ChatGPT-5 to 15.65 ± 3.36 for Perplexity Pro, and the mean FRES scores ranged from 37.50 ± 14.27 for Perplexity Pro to 52.59 ± 12.47 for Gemini 2.5.
conclusionCurrent LLMs may support oral cancer patient education, but their use remains limited by variable information quality, insufficient transparency, and poor readability. Although Perplexity Pro performed better on several reliability-related metrics, no model showed consistently high performance across all dimensions or met recommended readability standards. Future LLM-based patient education tools should prioritise verifiable sourcing, guideline-based accuracy, risk communication, and plain-language adaptation.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.