ArticleCureus2024
Assessing the Accuracy, Completeness, and Reliability of Artificial Intelligence-Generated Responses in Dentistry: A Pilot Study Evaluating the ChatGPT Model.
Article in Cureus, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 16 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
16 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Transformer-based models in dentistry: a systematic review.BMC medical informatics and decision making · 2026Pooled it
- Comparison of the diagnostic accuracy of dentists and ChatGPT in jawbone lesions.BMC oral health · 2026Article
- Evaluating the efficiency and factual reliability of LoRA for health misinformation detection.Scientific reports · 2026Article
- Generative Artificial Intelligence and Large Language Models in Paediatric Dentistry: A Scoping Review.International dental journal · 2026Article
- Comparative assessment of quality, consistency, and reference accuracy of MIH-related clinical information generated by ChatGPT-4o and DeepSeek R1.BMC oral health · 2026Article
- Generative Artificial Intelligence: Applications and Future Prospects in Dentistry.International dental journal · 2026Review
- Accuracy, readability, and bias of GPT-4o mini responses to oculoplastic patient questions.Frontiers in ophthalmology · 2026Article
- Assessing the suitability of ChatGPT in responding to public inquiries about dental crown restorations.BMC oral health · 2025Article
- The impact of language differences on the readability, quality, and reliability of information provided by artificial intelligence chatbots regarding vital pulp therapy: a cross-sectional study.BMC oral health · 2025Article
- Performance of AI-Chatbots to Common Temporomandibular Joint Disorders (TMDs) Patient Queries: Accuracy, Completeness, Reliability and Readability.Orthodontics & craniofacial research · 2025Article
- Comparative Evaluation of Responses from ChatGPT-5, Gemini 2.5 Flash, Grok 4, and Claude Sonnet-4 Chatbots to Questions About Endodontic Iatrogenic Events.Healthcare (Basel, Switzerland) · 2025Article
- Evaluation of ChatGPT-4's performance on pediatric dentistry questions: accuracy and completeness analysis.BMC oral health · 2025Article
- Assessing the Accuracy and Completeness of AI-Generated Dental Responses: An Evaluation of the Chat-GPT Model.Healthcare (Basel, Switzerland) · 2025Article
- Assessing ChatGPT's suitability in responding to the public's inquires on the effects of smoking on oral health.BMC oral health · 2025Article
- Comparing orthodontic pre-treatment information provided by large language models.BMC oral health · 2025Article
- The evaluation of tooth whitening from a perspective of artificial intelligence: a comparative analytical study.Frontiers in digital health · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundArtificial intelligence (AI) can be a tool in the diagnosis and acquisition of knowledge, particularly in dentistry, sparking debates on its application in clinical decision-making.
objectiveThis study aims to evaluate the accuracy, completeness, and reliability of the responses generated by Chatbot Generative Pre-Trained Transformer (ChatGPT) 3.5 in dentistry using expert-formulated questions. MATERIALS AND
methodsExperts were invited to create three questions, answers, and respective references according to specialized fields of activity. The Likert scale was used to evaluate agreement levels between experts and ChatGPT responses. Statistical analysis compared descriptive and binary question groups in terms of accuracy and completeness. Questions with low accuracy underwent re-evaluation, and subsequent responses were compared for improvement. The Wilcoxon test was utilized (α = 0.05).
resultsTen experts across six dental specialties generated 30 binary and descriptive dental questions and references. The accuracy score had a median of 5.50 and a mean of 4.17. For completeness, the median was 2.00 and the mean was 2.07. No difference was observed between descriptive and binary responses for accuracy and completeness. However, re-evaluated responses showed a significant improvement with a significant difference in accuracy (median 5.50 vs. 6.00; mean 4.17 vs. 4.80; p=0.042) and completeness (median 2.0 vs. 2.0; mean 2.07 vs. 2.30; p=0.011). References were more incorrect than correct, with no differences between descriptive and binary questions.
conclusionsChatGPT initially demonstrated good accuracy and completeness, which was further improved with machine learning (ML) over time. However, some inaccurate answers and references persisted. Human critical discernment continues to be essential to facing complex clinical cases and advancing theoretical knowledge and evidence-based practice.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.