ArticleBMC oral health2026
Comparison of artificial intelligence-based chatbots and expert periodontists in responding to patient questions: a multi-dimensional analysis.
Article in BMC oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundLarge Language Model (LLM) -based chatbots are increasingly used in patient information processes. The aim of this study was to compare the performance of ChatGPT (GPT-5.1), Gemini (2.5 Flash), and Claude (Sonnet 4.5) with expert periodontologists in responding to periodontal questions. Responses were evaluated in terms of scientific accuracy, completeness, conciseness & focus, empathy, and clarity, and differences among groups were investigated.
methodsThe question pool was developed de novo based on clinical experience. The literature and online search trends were used to support content coverage. The questions were evaluated using Lawshe's content validity method, and 20 open-ended questions were included. Expert responses were prepared by three experienced periodontologists based on the literature and clinical guidelines. The questions were then submitted using a standardized prompt to ChatGPT (GPT-5.1), Gemini (2.5 Flash), and Claude (Sonnet 4.5), and the first responses were recorded. All responses were anonymized and evaluated by nine independent periodontologists using a 5-point Likert scale in terms of scientific accuracy, completeness, conciseness & focus, empathy, and clarity. Statistical analyses were performed using the Friedman test with Bonferroni-corrected post-hoc comparisons, and inter-rater agreement was assessed using the intraclass correlation coefficient.
resultsNo significant difference was found among the groups in terms of scientific accuracy (p = 0.425). A significant difference was observed for completeness (p < 0.0001), with LLM-based chatbot responses achieving higher scores than expert responses. For conciseness & focus, a significant difference was found (p < 0.0001); Gemini demonstrated lower scores compared to the other groups, while no significant differences were observed among ChatGPT, Claude, and expert responses. A significant difference was observed for empathy (p < 0.0001), with all LLM-based chatbot responses scoring higher than expert responses. For clarity, a significant difference was found (p < 0.0001), with a difference observed only between ChatGPT and Gemini. Inter-rater agreement was within a good range across all domains.
conclusionsThis study showed that LLM-based chatbots can generate responses with a level of scientific accuracy comparable to expert periodontologists. However, despite providing more comprehensive and empathetic responses, some models demonstrated limitations in terms of conciseness & focus and clarity. These findings indicate that model differences should be considered in patient education and information-seeking processes. Expert responses were able to present a similar level of scientific accuracy using fewer and more focused expressions. This may make it more difficult for patients to maintain focus and to perceive information in a structured manner when using AI-generated responses. Overall, AI systems may serve a supportive role in patient information processes; however, physician supervision remains necessary in clinical use. The appropriate integration of LLM-based chatbots into patient information processes may enhance patient communication and access to information, provided that they are used as supportive tools rather than independent decision-makers in clinical practice.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.