Evidence map›Paper›PMID 41547810›Full record

ArticleBMC medical education2026

Comparative performance evaluation of ChatGPT-4 Omni and Gemini Advanced in the Turkish Dentistry Specialization Exam.

Makbule Buse Dundar Sari, Berkant Sezer

Abstract readComparative Study
In one paragraph

Article in BMC medical education, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Makbule Buse Dundar SariDepartment of Pediatric Dentistry, School of Dentistry, Çanakkale Onsekiz Mart University, Çanakkale, Türkiye.ORCID http://orcid.org/0000-0002-8848-8850
Berkant SezerDepartment of Pediatric Dentistry, School of Dentistry, Çanakkale Onsekiz Mart University, Çanakkale, Türkiye. dt.berkantsezer@gmail.com.ORCID http://orcid.org/0000-0001-9731-6156

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundIn recent years, advancements in artificial intelligence (AI) have led to the widespread integration of large language models and their chatbot applications into various fields, including dental education. This study aimed to evaluate the accuracy of ChatGPT-4 Omni (ChatGPT-4o) and Gemini Advanced in answering multiple-choice questions from the Turkish Dentistry Specialization Exams (DUS) across various disciplines.

methodsA total of 1,504 multiple-choice questions from 10 years of DUS exams were analyzed to compare the accuracy of ChatGPT-4o and Gemini Advanced. The questions were categorized into Fundamental Medical Sciences (n = 514) and Clinical Dental Sciences (n = 990). Each question was submitted to both chatbots, resulting in 3,008 responses. Accuracy was assessed using the official answer keys. Chi-square tests and Bonferroni post-hoc analyses were used to compare accuracy across disciplines and examine year-based variations.

resultsChatGPT-4o achieved an overall accuracy rate of 84%, while Gemini Advanced achieved 81.8% (p = 0.110). For the Fundamental Medical Sciences questions, no statistically significant differences were observed across sub-disciplines, with overall accuracies of 92.6% for ChatGPT-4o and 93.4% for Gemini Advanced. For the Clinical Dental Sciences questions, ChatGPT-4o outperformed Gemini Advanced in Prosthetic Dentistry (p = 0.013) and Dentomaxillofacial Radiology (p = 0.001), whereas Gemini Advanced showed higher accuracy in Pediatric Dentistry (p = 0.008). Across all Clinical Dental Sciences questions, ChatGPT-4o achieved an accuracy of 79.5%, compared to 75.8% for Gemini Advanced, and this difference was statistically significant (p = 0.046).

conclusionsAI-based chatbots demonstrate strong potential in answering multiple-choice dentistry questions. However, variations in performance across disciplines were observed, indicating differences in accuracy depending on the subject area. These findings highlight the potential educational implications of integrating AI into dental curricula, particularly as supplementary tools for exam preparation and knowledge reinforcement. Nevertheless, cautious integration is required to ensure that AI supports, rather than replaces, critical thinking and professional expertise.

Indexed as

Artificial IntelligenceEducational MeasurementEducation, DentalGenerative Artificial IntelligenceHumansLarge Language ModelsTurkeyArtificial intelligenceChatbotChatGPTDental educationDentistryExamGeminiLarge language modelsNatural language processing

Identifiers

PMID41547810
PMCPMC12895620

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.