Evidence map›Paper›PMID 41760912›Full record

ArticleNPJ digital medicine2026

A scalable framework for evaluating health language models.

Neil Mallinar, A Ali Heydari, Xin Liu, Anthony Z Faranesh, Brent Winslow, Nova Hammerquist, Benjamin Graef, Cathy Speed, Mark Malhotra, Shwetak Patel and 3 more

Abstract read
In one paragraph

Article in NPJ digital medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. A comprehensive survey of AI agents in healthcare.Journal of biomedical informatics · 2026
    Review
  3. Article
  4. Review
  5. Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Neil Mallinar *Google Research, Mountain View, CA, USA.
A Ali Heydari *Google Research, Mountain View, CA, USA.
Xin LiuGoogle Research, Mountain View, CA, USA.
Anthony Z FaraneshGoogle Research, Mountain View, CA, USA.
Brent WinslowGoogle Research, Mountain View, CA, USA.
Nova HammerquistGoogle Research, Mountain View, CA, USA.
Benjamin GraefVituity, Emeryville, CA, USA.
Cathy SpeedGoogle Research, Mountain View, CA, USA.
Mark MalhotraGoogle Research, Mountain View, CA, USA.
Shwetak PatelGoogle Research, Mountain View, CA, USA.
Javier L PrietoGoogle Research, Mountain View, CA, USA.
Daniel McDuffGoogle Research, Mountain View, CA, USA. dmcduff@google.com.
Ahmed A MetwallyGoogle Research, Mountain View, CA, USA. aametwally@google.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Large language models (LLMs) have emerged as powerful tools for analyzing and interpreting complex datasets. Recent studies demonstrate their potential to generate useful, personalized responses when provided with patient-specific health information that encompasses lifestyle, biomarkers, and context. As LLM-driven health applications are increasingly adopted, rigorous and efficient one-sided evaluation methodologies are crucial to ensure response quality across multiple dimensions, including accuracy, personalization, relevance and safety. However, current evaluation practices, particularly for open-ended text responses, heavily rely on human experts. This approach introduces human factors (perspectives, potential biases, inconsistencies) and is often cost-prohibitive, labor-intensive, and hinders scalability, especially in complex domains like healthcare where response assessment necessitates domain expertise and considers multifaceted patient data, which is often nuanced and diverse. In this work, we introduce Adaptive Precise Boolean rubrics: an evaluation framework that aims to streamline human and automated evaluation of open-ended questions by identifying critical gaps in model responses using a minimal set of targeted rubric questions. Our approach is based on recent work in more general evaluation settings that contrasts a smaller set of complex evaluation targets with a larger set of more precise, granular targets answerable with simple Boolean responses. We validate this approach in metabolic health, a domain encompassing diabetes, cardiovascular disease, and obesity. Our results demonstrate that Adaptive Precise Boolean rubrics yield substantially higher inter-rater agreement among both expert and non-expert human evaluators, as well as in automated assessments, compared to traditional Likert scales, while requiring approximately half the evaluation time of Likert-based methods. This enhanced efficiency and scalability, particularly through automated evaluation and non-expert contributions, paves the way for more extensive and cost-effective evaluation of LLMs in health.

Identifiers

PMID41760912
PMCPMC13249957

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.