ArticleCureus2026
Comparative Performance of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Oral and Maxillofacial Radiology Questions From the Turkish Dental Specialisation Examination.
Article in Cureus, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introductionRapid advances in artificial intelligence (AI) and large language models (LLMs) have increased interest in their use in dental education and assessment. This study evaluated the accuracy of GPT-5.1, Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1 on Turkish Dental Specialisation Exam (DUS) questions in oral and maxillofacial radiology.
methodsA total of 132 Turkish DUS questions from 2012 to 2021 were obtained from a publicly accessible question bank, including 126 theoretical and six image-based items. Each question had five options and one correct answer. Questions covered radiologic physics, radiation safety, dental anatomy, pathology classification, imaging instrumentation, and visual interpretation of panoramic, periapical, cone beam computed tomography (CBCT), and clinical images. All models received identical prompts, and image-based items were submitted through the image-upload functionality available in the respective web interfaces. Responses were scored against the official answer key by two oral and maxillofacial radiologists. Statistical analyses included Fisher's exact test, Cochran's Q test, and Bonferroni-corrected McNemar tests. An exploratory secondary analysis evaluated session-to-session variation using a model-based performance index.
resultsGPT-5.1 answered all 132 questions correctly (100%), compared with Claude Sonnet 4.5 (119/132; 90.2%), Gemini 2.0 Flash (111/132; 84.1%), and DeepSeek-R1 (108/132; 81.8%). Overall performance differed significantly among the models (p < 0.001), and GPT-5.1 significantly outperformed each of the other three models after Bonferroni correction. On the six image-based questions, accuracy was 100% for GPT-5.1, 83.3% for Claude Sonnet 4.5, 50.0% for Gemini 2.0 Flash, and 66.7% for DeepSeek-R1. Because only six image-based items were available, these subgroup findings were considered exploratory and do not permit firm conclusions regarding radiographic interpretation. No significant temporal trend was identified in the exploratory model-based performance index.
conclusionsGPT-5.1 achieved the highest accuracy on this publicly accessible DUS benchmark and significantly outperformed Claude Sonnet 4.5, Gemini 2.0 Flash, and DeepSeek-R1. However, the public availability of the questions and answer keys introduces a risk of benchmark contamination, and the very small image-based subgroup precludes generalisation regarding radiographic interpretation. These findings should therefore be interpreted as comparative performance on a public examination benchmark rather than evidence of clinical diagnostic competence.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.