Evidence map›Paper›PMID 42523813›Full record

ArticleFrontiers in public health2026

LLM-as-a-judge for infection prevention and control and antimicrobial resistance impact: comparing three main LLMs vs. human experts' assessment.

Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi

Abstract readComparative Study
In one paragraph

Article in Frontiers in public health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Marcello Di PumpoSection of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.
Leonardo VillaniSection of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.
Maria Rosaria GualanoUniCamillus-Saint Camillus International University of Health and Medical Sciences, Rome, Italy.
Danilo BuonsensoDepartment of Woman and Child Health, Fondazione Policlinico 'Agostino Gemelli' IRCCS, Rome, Italy.
Francesca RaffaelliDepartment of Laboratory and Infectivology Sciences, Fondazione Policlinico Universitario A. Gemelli IRCCS, Rome, Italy.
Daniele DonàDepartment of Women's and Children's Health, University of Padova, Padua, Italy.
Patrizia LaurentiSection of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.
Vittorio MaioJefferson College of Population Health, Thomas Jefferson University, Philadelphia, PA, United States.
Stefania BocciaSection of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.
Walter RicciardiSection of Hygiene, University Department of Life Science and Public Health, Università Cattolica del Sacro Cuore, Rome, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks. Methods: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines. Results: Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions. Conclusions: LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.

Indexed as

Drug Resistance, MicrobialInfection ControlLarge Language ModelsHumansReproducibility of Resultsantimicrobial resistance (AMR)large language models (LLM)LLM-as-a-judgeLLMs ratingpublic health

Identifiers

PMID42523813
PMCPMC13407514

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.