ReviewFrontiers in oral health2026
Multimodal large language models for oral lesion diagnosis: a systematic review of diagnostic performance and clinical utility.
Review in Frontiers in oral health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
14 citing papers in PubMed.
- Can artificial intelligence training improve clinical decision-making during deep caries excavation?Evidence-based dentistry · 2026Trial
- Limitations of off-the-shelf multimodal foundation models in secondary caries detection on interproximal radiographs.Odontology · 2026Article
- Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering Diseases: A Multicenter Paired Study.Diagnostics (Basel, Switzerland) · 2026Article
- Educational gaps and factors associated with artificial intelligence adoption among Egyptian periodontists: a multicenter cross-sectional study.Scientific reports · 2026Article
- Article
- Referral patterns to a Middle Eastern oral medicine service: a retrospective analysis.BMC health services research · 2026Article
- Guideline-based clinical reasoning in periodontology education: a comparative study of residents and large language models.BMC medical education · 2026Article
- A comparative analysis of large language models for providing oral cavity cancer information.Scientific reports · 2026Article
- Cognitive-level analysis of dentomaxillofacial radiology questions in the Turkish dentistry specialization examination: a Bloom's revised taxonomy analysis.BMC oral health · 2026Article
- MultiDentNet: a unified deep learning framework for multi-class dental condition screening and preliminary oral lesion triage.Scientific reports · 2026Article
- Abstraction-dependent diagnostic performance of a multimodal foundation model in oral epithelial dysplasia.Odontology · 2026Article
- Clinical and Patient Comparison of AI and Expert Digital Smile Design: A Prospective Paired Study.Dentistry journal · 2026Article
- Feasibility and exploratory assessment of large language models for pediatric dentistry queries: a comparative study.Frontiers in oral health · 2026Article
- Artificial intelligence for the detection and diagnosis of oral and maxillofacial lesions: evidence, limitations, and future directions.Brazilian oral research · 2026Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Diagnosing oral lesions from benign conditions to oral cancer remains challenging due to overlapping visual features and reliance on histopathology. Large language models (LLMs) can integrate textual and visual cues, but their diagnostic accuracy and clinical utility in real decision-making contexts remain uncertain. To systematically evaluate the diagnostic performance, clinical usefulness, and limitations of LLMs in identifying oral lesions. Methods: PubMed, CINAHL, Embase, Web of Science, and Google Scholar were searched to 20 July 2025. Eligible studies applied LLMs (e.g., ChatGPT, Gemini, DeepSeek, Copilot, Claude) for diagnosis or differential diagnosis of oral lesions using text, images, or multimodal inputs. Outcomes included diagnostic accuracy, agreement metrics, and qualitative assessments of explanation quality and clinical applicability. Risk of bias was assessed using an adapted QUADAS-2. Narrative synthesis was performed due to heterogeneity. Results: Seventeen studies (>1,200 cases) were included. Diagnostic accuracy ranged from 25%-96%, varying by model version, input modality, and lesion complexity. Multimodal inputs consistently improved performance, with Cohen's κ up to 0.85-0.90. Advanced models (GPT-4o, DeepSeek-R1, o1-preview) outperformed earlier versions and approached expert performance in some tasks, although specialists generally retained superior Top-1 accuracy. Clinical utility was highest when LLMs were used to structure differential reasoning, highlight red-flag features, and support communication, but limited in tasks requiring fine morphological interpretation or severity grading. Overall risk of bias was low to moderate. Conclusions: LLMs demonstrate variable diagnostic performance and context-dependent supportive utility as adjunctive tools in oral lesion assessment, particularly in multimodal settings. They should complement, rather than replace, expert clinical judgment. Future research should prioritize real-world workflow evaluation, standardized prompting strategies, and prospective clinical validation. Systematic Review Registration: https://www.crd.york.ac.uk/PROSPERO/view/CRD420251090315, identifier CRD420251090315.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.