ArticleBMC oral health2025
The impact of language differences on the readability, quality, and reliability of information provided by artificial intelligence chatbots regarding vital pulp therapy: a cross-sectional study.
Article in BMC oral health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
5 citing papers in PubMed.
- Can AI chatbots be reliable in dental emergencies? quality assessment of Arabic responses to dental emergency inquiries and public attitudes toward their use.BMC oral health · 2026Article
- Evaluation of Five Large Language Models for Parental Education in Pediatric Anesthesia: Reliability and Readability Study.JMIR medical informatics · 2026Article
- Large language models and generative artificial intelligence in endodontics: a scoping review.Odontology · 2026Review
- Evaluation of Arabic-Language AI Chatbot Responses to Migraine-Related Questions: A Comparative Cross-Sectional Study.Journal of clinical medicine · 2026Article
- Effect of language difference and time on the accuracy of artificial intelligence chatbots responses to questions about vital pulp therapy.BMC oral health · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe increasing use of artificial intelligence (AI) chatbots in healthcare has highlighted the need to evaluate the accuracy, reliability, and readability of the clinical information they provide. Vital pulp therapy is one of the fundamental biological approaches in modern dentistry aimed at preserving pulp vitality, and the quality of information related to this topic is highly important for clinical decision-making. The present study aimed to assess whether the readability, quality, and reliability of information provided by six different AI-based chatbots (ChatGPT, ChatGPT-4o, Gemini, Microsoft Copilot, Perplexity, and Claude) regarding vital pulp therapy vary depending on language differences.
methodsAfter a comprehensive literature review, 12 questions related to vital pulp therapy were developed. Each question was submitted to the chatbots in both Turkish and English for 7 consecutive days. The responses obtained in both languages were evaluated for readability using the Flesch Reading Ease Score (FRES) for English and the Ateşman Readability Formula for Turkish. Information quality was assessed using the Global Quality Scale (GQS), while reliability was evaluated based on the Journal of the American Medical Association (JAMA) benchmarks. Statistical analyses were performed using ANOVA, Bonferroni, and Chi-square tests, with a significance level set at p < 0.05.
resultsThe findings demonstrated significant language-based differences among the evaluated models. Readability, GQS, and JAMA assessments revealed statistically significant differences between the chatbots in both English and Turkish responses (p < 0.05). In the GQS evaluation, Gemini achieved the highest quality scores in English, while ChatGPT, ChatGPT-4o, and Claude produced the highest scores in Turkish (p < 0.05). In terms of JAMA reliability, Gemini and Perplexity showed the highest performance in English responses, whereas Perplexity demonstrated significantly higher reliability than the other platforms in Turkish (p < 0.05). Regarding readability, Perplexity generated the most difficult-to-read content in both languages, whereas ChatGPT provided the most readable responses (p < 0.05). The 7-day assessment showed no significant day-to-day changes in readability, quality, or reliability scores for most chatbots, indicating that their performance remained largely stable over time (p > 0.05).
conclusionsThis study demonstrated that language differences have a significant impact on the readability, quality, and reliability of information provided by AI–based chatbots regarding vital pulp therapy. These findings suggest that such systems may serve as supportive tools for accessing clinical information; however, expert oversight remains essential to ensure the accuracy and quality of the content. Future studies should include a wider variety of languages and chatbot models, along with extended evaluation periods.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.