Evidence map›Paper›PMID 40629684›Full record

ArticleMedical science monitor : international medical journal of experimental and clinical research2025

Performance of AI Chatbots in Preliminary Diagnosis of Maxillofacial Pathologies.

Ridvan Guler, Emine Yalcin

Abstract read
In one paragraph

Article in Medical science monitor : international medical journal of experimental and clinical research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Ridvan GulerDepartment of Oral and Maxillofacial Surgery, Dicle University Faculty of Dentistry, Diyarbakir, Turkey.ORCID 0000-0003-4750-9798
Emine YalcinDepartment of Oral and Maxillofacial Surgery, Dicle University Faculty of Dentistry, Diyarbakir, Turkey.ORCID 0009-0004-0978-8137

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

BACKGROUND Artificial intelligence (AI) has shown significant potential in transforming healthcare by enabling accurate, data-driven decision-making. This study compared the performance of the AI chatbots ChatGPT, Grok, Blackbox, and Claude AI in preliminary diagnosis of maxillofacial pathologies. MATERIAL AND METHODS This study included 23 patients (9 cysts, 14 neoplasms) who underwent operations at Dicle University Faculty of Dentistry between 2017 and 2024 and had their diagnoses histopathologically confirmed. For each case, 4 differential diagnosis options were prepared in question format and directed to the AI platforms. The accuracy of the answers given by the chatbots was analyzed by comparing them with the definitive histopathological diagnoses of the cases. Statistical analysis used the chi-square ad Fisher-Freeman-Halton tests to compare performance among the chatbots. Statistical significance was set at p<0.05. RESULTS ChatGPT answered 15 out of 23 questions correctly, achieving a success rate of 65.2%. Grok and Blackbox AI each achieved a success rate of 52.17%, while Claude AI achieved the lowest success rate, at 30.43%. When cases were categorized into cysts and neoplasms, Blackbox AI showed the highest accuracy for cyst cases (66.6%), while ChatGPT had the highest accuracy for neoplasm cases (71.4%). No statistically significant difference was observed in the distribution of correct and incorrect answers among the chatbots (p=0.125). No statistically significant difference was observed in the distribution of cysts and neoplasms answers among the chatbots (p=0.654). CONCLUSIONS Although all 4 AI chatbots achieved certain levels of accuracy, ChatGPT showed superior performance compared to other chatbots. The development of these chatbots could be beneficial for diagnostic accuracy and treatment recommendations in dentistry.

Indexed as

Artificial IntelligenceAdultCystsDiagnosis, DifferentialFemaleGenerative Artificial IntelligenceHumansMaleMiddle Aged

Identifiers

PMID40629684
PMCPMC12257980

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.