ArticleOdontology2026
Clinical safety, content coverage, and patient-centered language of AI responses to periodontal complaint-based queries: a comparative study of four large language models.
Article in Odontology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
This study aimed to evaluate the clinical safety, informational completeness, and patient-centered language of responses generated by four large language model (LLM)-based systems to standardized periodontal complaint-based queries. Seven standardized symptom-based periodontal queries were developed based on commonly reported patient complaints and submitted in Turkish to Copilot, Gemini, Claude, and ChatGPT. Responses were evaluated using a structured rule-based framework consisting of a Content Coverage Score (CCS), Risk of Harm Score (RHS), and Patient Language Score (PLS). Assessments were performed by two blinded human reviewers and one blinded AI-based evaluator. Human consensus, AI consensus, and combined evaluator scores were calculated. Agreement between evaluators was assessed using intraclass correlation coefficients (ICC), while human-AI differences and inter-model comparisons were analyzed using non-parametric statistical tests. Excellent agreement was observed between human reviewers, repeated AI evaluations, and human-AI consensus scores (ICC range: 0.911-0.925; p < 0.001). No significant difference was found between overall human and AI consensus scores (p = 0.985). Across the four AI systems, no statistically significant differences were observed in CCS or PLS scores in the human, AI, or combined evaluator analyses (all p > 0.05). PLS scores were generally high across models, indicating good linguistic accessibility for patients. No clearly harmful guidance was identified by the human reviewers. Overall, LLM-based systems generated clinically safe, reasonably comprehensive, and generally patient-accessible responses to common periodontal complaint-based queries. Although these systems may serve as supplementary sources of periodontal health information, they cannot replace individualized clinical evaluation and professional dental consultation.
Indexed as
Identifiers
42412385What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.