Evidence map›Paper›PMID 42761307›Full record

ArticleCureus2026

Comparative Performance of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Oral and Maxillofacial Radiology Questions From the Turkish Dental Specialisation Examination.

Gaye Keser, Burcu Celen, Filiz Namdar Peki̇ner

Abstract read
In one paragraph

Article in Cureus, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Gaye KeserDepartment of Oral Medicine and Radiology, Marmara University, Istanbul, TUR.
Burcu CelenDepartment of Oral Medicine and Radiology, Marmara University, Istanbul, TUR.
Filiz Namdar Peki̇nerDepartment of Oral and Maxillofacial Radiology, Marmara University, Istanbul, TUR.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionRapid advances in artificial intelligence (AI) and large language models (LLMs) have increased interest in their use in dental education and assessment. This study evaluated the accuracy of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Turkish Dental Specialisation Exam (DUS) questions in oral and maxillofacial radiology.

methodsA total of 132 Turkish DUS questions from 2012 to 2021 were obtained from a publicly accessible question bank, including 126 theoretical and six image-based items. Each question had five options and one correct answer. Questions covered radiologic physics, radiation safety, dental anatomy, pathology classification, imaging instrumentation, and visual interpretation of panoramic, periapical, cone beam computed tomography (CBCT), and clinical images. All models received identical prompts, and image-based items were submitted through the image-upload functionality available in the respective web interfaces. Responses were scored against the official answer key by two oral and maxillofacial radiologists. Statistical analyses included Fisher's exact test, Cochran's Q test, and Bonferroni-corrected McNemar tests. An exploratory secondary analysis evaluated session-to-session variation using a model-based performance index.

resultsGPT-5.1 answered all 132 questions correctly (100%), compared with Claude Sonnet 4.5 (119/132; 90.2%), Gemini 2.0 Flash (111/132; 84.1%), and DeepSeek-R1 (108/132; 81.8%). Overall performance differed significantly among the models (p < 0.001), and GPT-5.1 significantly outperformed each of the other three models after Bonferroni correction. On the six image-based questions, accuracy was 100% for GPT-5.1, 83.3% for Claude Sonnet 4.5, 50.0% for Gemini 2.0 Flash, and 66.7% for DeepSeek-R1. Because only six image-based items were available, these subgroup findings were considered exploratory and do not permit firm conclusions regarding radiographic interpretation. No significant temporal trend was identified in the exploratory model-based performance index.

conclusionsGPT-5.1 achieved the highest accuracy on this publicly accessible DUS benchmark and significantly outperformed Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1. However, the public availability of the questions and answer keys introduces a risk of benchmark contamination, and the very small image-based subgroup precludes generalisation regarding radiographic interpretation. These findings should therefore be interpreted as comparative performance on a public examination benchmark rather than evidence of clinical diagnostic competence.

Indexed as

board examdeep learning artificial intelligencelarge language models (llms)specialty doctorstudent education

Identifiers

PMID42761307
PMCPMC13586900

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.