Evidence map›Paper›PMID 41408270›Full record

ArticleBMC oral health2025

The impact of language differences on the readability, quality, and reliability of information provided by artificial intelligence chatbots regarding vital pulp therapy: a cross-sectional study.

Emine Şimşek, Özge Kurt

Abstract read
In one paragraph

Article in BMC oral health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Article
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Emine ŞimşekDepartment of Endodontics, Faculty of Dentistry, Mersin University, Mersin, Turkey. eminesimsek@mersin.edu.tr.ORCID http://orcid.org/0000-0001-9195-2012
Özge KurtDepartment of Endodontics, Faculty of Dentistry, Aksaray University, Aksaray, Turkey.ORCID http://orcid.org/0000-0001-6261-1200

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe increasing use of artificial intelligence (AI) chatbots in healthcare has highlighted the need to evaluate the accuracy, reliability, and readability of the clinical information they provide. Vital pulp therapy is one of the fundamental biological approaches in modern dentistry aimed at preserving pulp vitality, and the quality of information related to this topic is highly important for clinical decision-making. The present study aimed to assess whether the readability, quality, and reliability of information provided by six different AI-based chatbots (ChatGPT, ChatGPT-4o, Gemini, Microsoft Copilot, Perplexity, and Claude) regarding vital pulp therapy vary depending on language differences.

methodsAfter a comprehensive literature review, 12 questions related to vital pulp therapy were developed. Each question was submitted to the chatbots in both Turkish and English for 7 consecutive days. The responses obtained in both languages were evaluated for readability using the Flesch Reading Ease Score (FRES) for English and the Ateşman Readability Formula for Turkish. Information quality was assessed using the Global Quality Scale (GQS), while reliability was evaluated based on the Journal of the American Medical Association (JAMA) benchmarks. Statistical analyses were performed using ANOVA, Bonferroni, and Chi-square tests, with a significance level set at p < 0.05.

resultsThe findings demonstrated significant language-based differences among the evaluated models. Readability, GQS, and JAMA assessments revealed statistically significant differences between the chatbots in both English and Turkish responses (p < 0.05). In the GQS evaluation, Gemini achieved the highest quality scores in English, while ChatGPT, ChatGPT-4o, and Claude produced the highest scores in Turkish (p < 0.05). In terms of JAMA reliability, Gemini and Perplexity showed the highest performance in English responses, whereas Perplexity demonstrated significantly higher reliability than the other platforms in Turkish (p < 0.05). Regarding readability, Perplexity generated the most difficult-to-read content in both languages, whereas ChatGPT provided the most readable responses (p < 0.05). The 7-day assessment showed no significant day-to-day changes in readability, quality, or reliability scores for most chatbots, indicating that their performance remained largely stable over time (p > 0.05).

conclusionsThis study demonstrated that language differences have a significant impact on the readability, quality, and reliability of information provided by AI–based chatbots regarding vital pulp therapy. These findings suggest that such systems may serve as supportive tools for accessing clinical information; however, expert oversight remains essential to ensure the accuracy and quality of the content. Future studies should include a wider variety of languages and chatbot models, along with extended evaluation periods.

Indexed as

Artificial IntelligenceComprehensionLanguageCross-Sectional StudiesHumansReproducibility of ResultsTurkeyArtificial intelligenceChatbotVital pulp therapy

Identifiers

PMID41408270
PMCPMC12821960

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.