Evidence map›Paper›PMID 41299627›Full record

ArticleBMC medical education2025

Knowledge-level comparison in pulpal and periapical diseases: dental students versus artificial intelligence models (Gemini, Microsoft Copilot, ChatGPT-3.5, ChatGPT-4o): cross-sectional study.

Özge Kurt, Emine Şimsek

Abstract readComparative Study
In one paragraph

Article in BMC medical education, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Özge KurtFaculty of Dentistry, Department of Endodontics, Aksaray University, Bahçesaray Neighborhood, Necmettin Erbakan Boulevard, Campus Road, Aksaray, 68100, Turkey. Ozgekrtrs@gmail.com.ORCID http://orcid.org/0000-0001-6261-1200
Emine ŞimsekFaculty of Dentistry, Department of Endodontics, Mersin University, Mersin, Turkey.ORCID http://orcid.org/0000-0001-9195-2012

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThis study explored the diagnostic accuracy of artificial intelligence (AI) chatbots and dental students when responding to questions related to pulpal and periapical diseases. Rapid advancements in AI have led to increased interest in their applicability to clinical education and decision-making in dentistry.

objectiveTo compare the accuracy rates of responses given by dental students and various AI-based chatbots (ChatGPT-3.5, ChatGPT-4o, Gemini, and Microsoft Copilot) to multiple-choice questions designed to assess knowledge related to pulpal and periapical diseases.

methodsThe study included third- and fifth-year dental students representing different levels of clinical training, along with four distinct AI-based chatbots. A total of 327 responses were collected from students, while each chatbot generated 450 responses. The evaluation was based on 15 multiple-choice questions developed in accordance with the 2020 version of the American Association of Endodontists (AAE) clinical guidelines. The accuracy rates of the groups were compared using descriptive statistics, one-way ANOVA, Bonferroni post hoc tests for significant differences, and Chi-square tests for correct versus incorrect response ratios.

resultsThe highest accuracy rate was observed among fifth-year dental students (85.1%), followed by ChatGPT-4o (79.6%), ChatGPT-3.5 (75.1%), Gemini (71.6%), third-year students (64.9%), and Microsoft Copilot (61.3%). A statistically significant difference was found among the groups (p < 0.05). ChatGPT-4o demonstrated a comparable accuracy rate to fifth-year students with more clinical experience (p > 0.05), whereas other chatbots and third-year students showed lower performance.

conclusionChatbots exhibited varying levels of accuracy in diagnosing pulpal and periapical diseases. ChatGPT-4o performed at a level similar to that of more clinically experienced students, suggesting its potential as a supportive tool in dental education and clinical decision support systems. However, the relatively lower accuracy rates of models such as Gemini and Microsoft Copilot underscore the continued importance of human expertise. These findings suggest that while AI systems may serve as complementary tools in education, they cannot fully replace clinical judgment grounded in human experience.

Indexed as

Artificial IntelligenceDental Pulp DiseasesEducation, DentalPeriapical DiseasesStudents, DentalClinical CompetenceCross-Sectional StudiesEducational MeasurementFemaleGenerative Artificial IntelligenceHumansMaleArtificial intelligenceDentalDental studentsEducationLarge language models periapical periodontitisPulp disease

Identifiers

PMID41299627
PMCPMC12659035

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.