Evidence map›Paper›PMID 42087133›Full record

Observational studyBMC oral health2026

Evaluating the clinical safety of large language models in oral cancer-related patient communication: a repeated-prompt observational study.

Burcu Yeliz Kollayan, Tuğba Cebeci

Abstract readObservational Study
In one paragraph

Observational study in BMC oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Burcu Yeliz KollayanDepartment of Oral and Maxillofacial Radiology, Faculty of Dentistry, Mehmet Akif Ersoy University, Burdur, Turkey. byevran@hotmail.com.ORCID 0000-0001-9609-7019
Tuğba CebeciDepartment of Oral and Maxillofacial Radiology, Faculty of Dentistry, İstanbul Aydın University, Istanbul, Turkey.ORCID 0009-0009-8068-4652

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAs patients increasingly consult large language models (LLMs) for health-related information, evaluating the clinical safety of AI-generated responses has become essential, particularly in high-risk domains such as oral oncology. Despite growing interest in AI applications in medicine, evidence regarding response consistency and referral safety in patient communication remains limited. This study aimed to assess the clinical safety of two contemporary LLMs in oral cancer-related patient scenarios using a multidimensional evaluation framework.

methodsThis repeated-prompt observational study evaluated Google Gemini (Pro version) and xAI Grok (Grok-1) over a 7-day period. Twenty standardized Turkish-language patient scenarios related to suspected oral cancer were submitted daily to each model, generating 280 responses. Scientific accuracy and completeness were assessed using a 5-point Likert scale by two independent oral and maxillofacial radiologists. Readability was evaluated using validated Turkish indices (Ateşman and Bezirci-Yılmaz). Referral safety was assessed as a binary outcome. Internal consistency across repeated prompts was measured using Cronbach's alpha, and inter-model agreement was analyzed using intraclass correlation coefficients (ICC).

resultsBoth models demonstrated comparable levels of scientific accuracy (Gemini: 3.52 ± 0.57; Grok: 3.39 ± 0.68; p = 0.072) and completeness (3.40 ± 0.70 vs. 3.25 ± 0.78; p = 0.091). Overall referral safety was high (Gemini: 90.0%; Grok: 92.1%), although the Gemini model failed to recommend professional consultation in two high-risk scenarios involving suspected malignancy. In contrast, Grok consistently recommended referral across all scenarios. Readability scores were similar between models; however, Grok generated significantly longer sentences (p = 0.0005; Cohen's d = 2.50), indicating increased linguistic complexity. Internal consistency was high for both models (Gemini α = 0.942; Grok α = 0.886), whereas inter-model agreement was moderate (ICC: 0.50-0.58).

conclusionsContemporary LLMs demonstrate generally acceptable accuracy and a precautionary approach in oral cancer-related communication. However, variability in referral behavior and linguistic structure, along with occasional under-referral in high-risk scenarios, highlights potential clinical risks. While these systems may support patient education and triage, they should be considered adjunctive tools and not substitutes for professional evaluation. Further research is needed to optimize the balance between clinical caution and appropriate guidance in AI-assisted healthcare communication.

Indexed as

CommunicationLarge Language ModelsMouth NeoplasmsPatient SafetyComprehensionFemaleHumansMaleArtificial intelligenceLarge language modelsOral cancerPatient communicationReadability

Identifiers

PMID42087133
PMCPMC13289552

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.