Evidence map›Paper›PMID 42093540›Full record

ArticleAnnals of medicine2026

Assessing diagnostic performance of multimodal AI and human experts in oral and maxillofacial radiography: a comparative analysis of ChatGPT, Grok, and MANUS.

Ahmed A Madfa, Abdullah F Alshammari, Bassam A Anazi

Abstract readComparative Study
In one paragraph

Article in Annals of medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Ahmed A MadfaDepartment of Restorative Dental Science, College of Dentistry, University of Ha'il, Ha'il, Kingdom of Saudi Arabia.ORCID 0000-0001-6124-0129
Abdullah F AlshammariDepartment of Basic Dental and Medical Science, College of Dentistry, University of Ha'il, Ha'il, Kingdom of Saudi Arabia.ORCID 0000-0002-5200-6362
Bassam A AnaziDepartment of Basic Dental and Medical Science, College of Dentistry, University of Ha'il, Ha'il, Kingdom of Saudi Arabia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundArtificial intelligence (AI), particularly large language models (LLMs), is increasingly applied to radiographic interpretation in healthcare. In dentistry, radiographic imaging is essential for diagnosis and treatment planning, yet remains subject to variability and human error. AI may enhance diagnostic accuracy and consistency.

aimTo evaluate and compare the diagnostic accuracy, consistency, and interpretive performance of multimodal AI models-ChatGPT, Grok, and MANUS-with expert radiologists in dental radiograph interpretation.

methodsA total of 120 anonymised radiographs (40 OPGs, 40 periapical, 40 CT slices) were selected from validated academic sources. Two board-certified oral and maxillofacial radiologists established gold standard diagnoses. Each image was independently assessed by the three AI models under standardised prompting. Diagnostic accuracy and intra-model consistency were analysed using descriptive statistics, Cohen's kappa, McNemar's test, and logistic regression.

resultsIn the first assessment, MANUS and ChatGPT achieved 92.5% accuracy (111/120), while Grok reached 88.3% (106/120). Performance improved in the second round: MANUS 95.0%, ChatGPT 93.3%, and Grok 90.8%, compared with 96.7% for radiologists. ChatGPT showed the highest reproducibility (κ = 0.937), whereas MANUS demonstrated the highest overall accuracy. Strong agreement was observed between ChatGPT and MANUS, with greater variability in Grok. No significant systematic bias was detected between AI outputs and radiologist benchmarks.

conclusionThe evaluated LLMs demonstrated diagnostic performance comparable to expert radiologists. MANUS excelled in accuracy and ChatGPT in reproducibility, supporting their potential as adjunct tools in dental radiology, while maintaining the need for expert oversight.Clinical trial number: Not applicable.

Indexed as

Artificial IntelligenceRadiographic Image Interpretation, Computer-AssistedRadiography, DentalFemaleGenerative Artificial IntelligenceHumansLarge Language ModelsRadiologistsReproducibility of ResultsTomography, X-Ray ComputedArtificial intelligenceChatGPTdentistryGroklarge language modelsMANUSradiographic diagnostics

Identifiers

PMID42093540
PMCPMC13159575

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.