Evidence map›Paper›PMID 42780134›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Validating LLM judges for automated oversight of patient communication.

Zidu Xu, Johnathan Zeng, Shuang Zhou, Zhihong Zhang, Thibault Heintz, Marion Tonneau, Arsalan Yaghoubi, Bingyang Ye, Vikram Goddla, Lisa Lehmann and 10 more

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

20 authors.

Zidu XuArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.ORCID 0000-0002-6122-8426
Johnathan ZengArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.ORCID 0000-0002-0838-9263
Shuang ZhouDepartment of Neurology, Massachusetts General Hospital, Boston, MA, USA.
Zhihong ZhangColumbia University School of Nursing, Columbia University, New York, NY, USA.
Thibault HeintzArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.
Marion TonneauArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.
Arsalan YaghoubiDepartment of Computer Science, Loyola University Chicago, Chicago, IL, USA.
Bingyang YeArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.ORCID 0009-0001-1696-4248
Vikram GoddlaHarvard College, Harvard University, Cambridge, MA, USA.ORCID 0000-0001-6286-6684
Lisa LehmannDepartment of Medicine, Mass General Brigham, Harvard Medical School, Boston, MA, USA.
Yu-Hui ChenDepartment of Data Science, Dana-Farber Cancer Institute, Boston, MA, USA.
Elad SharonDepartment of Medical Oncology, Dana-Farber Cancer Institute, Boston, MA, USA.
David E KozonoDepartment of Radiation Oncology, Brigham and Women's Hospital/Dana-Farber Cancer Institute, Boston, MA, USA.
Anna RevetteDivision of Population Sciences, Dana-Farber Cancer Institute, Boston, MA, USA.
Julia MauesGuiding Researchers and Advocates to Scientific Partnerships (GRASP), Baltimore, MD, USA.
Thelma BrownNCI Breast Cancer Steering Committee, National Cancer Institute, Bethesda, MD, USA.
Paul CatalanoDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
Raymond H MakArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.
Dimitry DligachDepartment of Computer Science, Loyola University Chicago, Chicago, IL, USA.
Danielle S BittermanArtificial Intelligence in Medicine Program, Mass General Brigham, Harvard Medical School, Boston, MA, USA.ORCID 0000-0003-0345-2232

Funding

Harvard Clinical and Translational Science CenterUM1TR004408 · NCATS · HARVARD MEDICAL SCHOOL · PI Lindsey Robert Baden, Lee Marshall Nadler · 2023 to 2026
$43.3M
Shared Resource Core 2: Clinical Artificial Intelligence CoreU54CA274516 · NCI · DANA-FARBER CANCER INST · PI Ross I. Berbeco · 2023 to 2026
$8.1M
Cancer Deep Phenotyping from Electronic Medical RecordsU24CA248010 · NCI · BOSTON CHILDREN'S HOSPITAL · PI HARRY S HOCHHEISER, GUERGANA K. SAVOVA · 2020 to 2026
$6.1M
Informatics strategies to improve immune-related adverse event detection in cancer patientsR01CA294033 · NCI · BRIGHAM AND WOMEN'S HOSPITAL · PI Hugo Aerts · 2024 to 2026
$1.3M
NCATS NIH HHS UM1 TR004408NCI NIH HHS R01 CA294033NCI NIH HHS U24 CA248010NCI NIH HHS U54 CA274516
6 · The paper itself

Abstract

LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints. We developed a stakeholder-informed seven-criterion evaluation rubric spanning safety, reliability, and communication quality. Clinician reference ratings showed strong interrater reliability across all criteria. We validated the rubric on the Informed CONsent Benchmark (ICON-Bench) and benchmarked 19 LLM judges across multiple implementation strategies. LLM judges achieved strong clinician agreement for safety screening and factual verification (Spearman

Identifiers

PMID42780134
PMCPMC13596599

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.