Evidence map›Paper›PMID 41816099›Full record

ReviewFrontiers in oral health2026

Multimodal large language models for oral lesion diagnosis: a systematic review of diagnostic performance and clinical utility.

Fatma E A Hassanein, Malik Alkabazi, Melek Tassoker, Yousra Ahmed, Suliman Alsaeed, Asmaa Abou-Bakr

Abstract readReview
In one paragraph

Review in Frontiers in oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.

0numbers the graph read from it
0cells of the map it votes in
14citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

14 citing papers in PubMed.

  1. Trial
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Fatma E A HassaneinDepartment of Oral Medicine, Periodontology, and Oral Diagnosis, Faculty of Dentistry, King Salman International University, El Tur, South Sinai, Egypt.
Malik AlkabaziFaculty of Dentistry Khalij-Libya, Tripoli, Libya.
Melek TassokerDepartment of Dentomaxillofacial Radiology, Faculty of Dentistry, Necmettin Erbakan University, Meram, Konya, Türkiye.
Yousra AhmedDepartment of Prosthetic Dentistry, Removable Prosthodontic Division, Faculty of Dentistry, King Salman International University, El Tur, South Sinai, Egypt.
Suliman AlsaeedPreventive Dental Sciences Department, College of Dentistry, King Saud Bin Abdulaziz University for Health Sciences, Riyadh, Saudi Arabia.
Asmaa Abou-BakrDepartment of Oral Medicine and Periodontology, Faculty of Dentistry, Galala University, Suez, Egypt.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Diagnosing oral lesions from benign conditions to oral cancer remains challenging due to overlapping visual features and reliance on histopathology. Large language models (LLMs) can integrate textual and visual cues, but their diagnostic accuracy and clinical utility in real decision-making contexts remain uncertain. To systematically evaluate the diagnostic performance, clinical usefulness, and limitations of LLMs in identifying oral lesions. Methods: PubMed, CINAHL, Embase, Web of Science, and Google Scholar were searched to 20 July 2025. Eligible studies applied LLMs (e.g., ChatGPT, Gemini, DeepSeek, Copilot, Claude) for diagnosis or differential diagnosis of oral lesions using text, images, or multimodal inputs. Outcomes included diagnostic accuracy, agreement metrics, and qualitative assessments of explanation quality and clinical applicability. Risk of bias was assessed using an adapted QUADAS-2. Narrative synthesis was performed due to heterogeneity. Results: Seventeen studies (>1,200 cases) were included. Diagnostic accuracy ranged from 25%-96%, varying by model version, input modality, and lesion complexity. Multimodal inputs consistently improved performance, with Cohen's κ up to 0.85-0.90. Advanced models (GPT-4o, DeepSeek-R1, o1-preview) outperformed earlier versions and approached expert performance in some tasks, although specialists generally retained superior Top-1 accuracy. Clinical utility was highest when LLMs were used to structure differential reasoning, highlight red-flag features, and support communication, but limited in tasks requiring fine morphological interpretation or severity grading. Overall risk of bias was low to moderate. Conclusions: LLMs demonstrate variable diagnostic performance and context-dependent supportive utility as adjunctive tools in oral lesion assessment, particularly in multimodal settings. They should complement, rather than replace, expert clinical judgment. Future research should prioritize real-world workflow evaluation, standardized prompting strategies, and prospective clinical validation. Systematic Review Registration: https://www.crd.york.ac.uk/PROSPERO/view/CRD420251090315, identifier CRD420251090315.

Indexed as

artificial intelligence in dentistryclinical decision supportdiagnostic accuracylarge language modelsmultimodal AIoral lesions

Identifiers

PMID41816099
PMCPMC12971682

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.