Evidence map›Paper›PMID 42150339›Full record

SynthesisInternational dental journal2026

Accuracy of Large Language Models in Answering Dental Examination Questions: A Systematic Review and Meta-Analysis.

Mahmood Dashti, Farshad Khosraviani, Atieh Meyari, Mohammad Hosein Amirzade-Iranaq, Akhilanand Chaurasia, Delband Hefzi, Niloofar Ghadimi, Antonin Tichy, Zohaib Khurshid, Falk Schwendicke

Abstract readSystematic ReviewMeta-Analysis
In one paragraph

Synthesis in International dental journal, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Mahmood DashtiDentofacial Deformities Research Center, Research Institute of Dental Sciences, Shahid Beheshti University of Medical Sciences, Tehran, Iran; Department of Artificial Intelligence Engineering, Graduate School of Natural and Applied Sciences, Istinye University, Istanbul, Türkiye. Electronic address: 2533275099@stu.istinye.edu.tr.
Farshad KhosravianiUCLA School of Dentistry, CA, USA.
Atieh MeyariDepartment of Esthetic & Restorative Dentistry, School of Dentistry, Tehran Azad University of Medical Sciences, Tehran, Iran.
Mohammad Hosein Amirzade-IranaqUniversal Scientific Education and Research Network (USERN), Tehran University of Medical Sciences, Tehran, Iran.
Akhilanand ChaurasiaDepartment of Oral Medicine and Radiology, King George's Medical University, Lucknow, India.
Delband HefziSchool of Dentistry, Tehran University of Medical Science, Tehran, Iran.
Niloofar GhadimiDental Materials Research Center, TeMS.C., School of Dentistry, Azad University of Medical Sciences, Tehran, Iran.
Antonin TichyClinic for Conservative Dentistry, Periodontology and Digital Dentistry, LMU University Hospital, LMU Munich, Munich, Germany; Institute of Dental Medicine, First Faculty of Medicine, Charles University, Prague, Czech Republic.
Zohaib KhurshidDepartment of Prosthodontics and Dental Implantology, College of Dentistry, King Faisal University, Al-Ahsa, Saudi Arabia; Center for Artificial Intelligence and Innovation (CAII), Faculty of Dentistry, Chulalongkorn University, Bangkok, Thailand.
Falk SchwendickeClinic for Conservative Dentistry, Periodontology and Digital Dentistry, LMU University Hospital, LMU Munich, Munich, Germany. Electronic address: falk.schwendicke@med.uni-muenchen.de.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionLarge language models (LLMs), including OpenAI's GPT family accessed via interfaces such as ChatGPT and Microsoft Copilot, as well as non-GPT systems such as Google Gemini, are increasingly applied in healthcare and dental education. However, the accuracy of these systems in specialized tasks such as answering dental examination questions remains unclear.

methodsThis systematic review and meta-analysis evaluated LLM performance in answering dental questions. Databases searched were PubMed, Embase, Scopus, and Web of Science. Data on question type and number, LLM versions, and accuracy rates were extracted. Pooled accuracy was estimated using a random-effects model; heterogeneity and publication bias were assessed.

resultsA total of 39 studies were included, with ChatGPT-4 being the most frequently evaluated model. The pooled accuracy for LLMs was 63.7% (95% CI: 60.3%-67.1%), with high heterogeneity (I² = 91.5%). Subgroup analysis revealed ChatGPT-4 and Copilot (a GPT-based interface) achieved the highest pooled accuracies (∼73% and ∼75%, respectively). Direct comparisons confirmed ChatGPT-4 significantly outperformed earlier versions and some competitor models. Sensitivity analyses supported the robustness of findings.

conclusionLLMs demonstrate moderate accuracy in answering dental examination questions and are currently insufficient for autonomous clinical decision-making. When their limitations are explicitly recognized, however, these systems may serve as valuable adjuncts in dental education and examination preparation. Methodological strategies such as structured prompting and retrieval-augmented approaches warrant further investigation but were not the primary focus of the present analysis.

Indexed as

Educational MeasurementEducation, DentalLarge Language ModelsHumansArtificial intelligenceDental examinationLarge language modelsmeta-analysis

Identifiers

PMID42150339
PMCPMC13202568

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.