Evidence map›Paper›PMID 42402252›Full record

ArticleJournal of clinical periodontology2026

Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A Blinded Expert-Rated Comparative Study of 11 Systems.

Yaniv Mayer, Bertha Demetriou, Giulio Rasperini, Eran Gabay, Rok Gašperšič, Darko Božić, Javier Calatrava, Ofir Ginesin, Hadar Zigdon Giladi, Ricardo Faria Almeida

Abstract readComparative Study
In one paragraph

Article in Journal of clinical periodontology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Yaniv MayerThe Ruth and Bruce Rappaport Faculty of Medicine, Technion, Israel Institute of Technology, Haifa, Israel.ORCID https://orcid.org/0000-0001-5500-7961
Bertha DemetriouDepartment of Periodontology, Rambam Health Care Campus, Haifa, Israel.
Giulio RasperiniDepartment of Biomedical, Surgical and Dental Sciences, University of Milan, Milan, Italy.ORCID https://orcid.org/0000-0003-3836-147X
Eran GabayThe Ruth and Bruce Rappaport Faculty of Medicine, Technion, Israel Institute of Technology, Haifa, Israel.ORCID https://orcid.org/0000-0003-0429-2883
Rok GašperšičDepartment of Oral Medicine and Periodontology, Faculty of Medicine, University of Ljubljana, Ljubljana, Slovenia.
Darko BožićDepartment of Periodontology, School of Dental Medicine, University of Zagreb, Zagreb, Croatia.ORCID https://orcid.org/0000-0002-0391-0574
Javier CalatravaSection of Graduate Periodontology, Faculty of Odontology, University Complutense of Madrid, Madrid, Spain.ORCID https://orcid.org/0009-0007-1080-6285
Ofir GinesinThe Ruth and Bruce Rappaport Faculty of Medicine, Technion, Israel Institute of Technology, Haifa, Israel.ORCID https://orcid.org/0000-0002-8736-9753
Hadar Zigdon GiladiThe Ruth and Bruce Rappaport Faculty of Medicine, Technion, Israel Institute of Technology, Haifa, Israel.
Ricardo Faria AlmeidaLAQV/REQUIMTE, Periodontology, Oral Surgery and Oral Medicine Department, Faculty of Dental Medicine, University of Porto, Porto, Portugal.ORCID https://orcid.org/0000-0001-9265-0913

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

aimTo compare retrieval-augmented systems with general-purpose large language models (LLMs) on standardised periodontal clinical vignettes. MATERIALS AND

methodsEleven AI systems were evaluated: nine general-purpose LLMs, one general-purpose retrieval-augmented platform (Perplexity) and one medical-domain retrieval-augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5-point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover-Iman tests with Holm correction were applied; mixed-effects and ordinal models served as sensitivity analyses.

resultsAt least one parameter scored dangerous (≤ 2) in 3.3%-46.7% of responses across platforms, despite mean composite scores (3.28-4.86) exceeding the rubric midpoint of 3.0. Between-model differences were significant (p < 0.001), with a small-to-medium overall effect (Kendall's W = 0.17) and large within-category effects (W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier.

conclusionRetrieval-augmented systems rated highest, but this advantage was confounded with response length. The dangerous-response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.

Indexed as

Information Storage and RetrievalLarge Language ModelsPeriodontal DiseasesGenerative Artificial IntelligenceHumansartificial intelligenceclinical decision supportdiagnosislarge language modelsperiodontitis

Identifiers

PMID42402252
PMCPMC13581464

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.