Evidence map›Paper›PMID 39405325›Full record

SynthesisJAMA2025

Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review.

Suhana Bedi, Yutong Liu, Lucy Orr-Ewing, Dev Dash, Sanmi Koyejo, Alison Callahan, Jason A Fries, Michael Wornow, Akshay Swaminathan, Lisa Soleymani Lehmann and 9 more

Registry-linked trialAbstract readSystematic Review
In one paragraph

Synthesis in JAMA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT06997107 (AI-CARE), which is not on this map. Cited by 357 papers, 5 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
357citing papers in PubMed, 5 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT06997107 naactive not recruitingnot on this map

AI-CARE: AI to Create Accessible & Reliable Patient Education Materials

TypeinterventionalSponsorHopital MontfortRan2025 to 2026Enrolled50ConditionsPrimary Care PatientsArmsHealth promotional messages generated by Artificial Intelligence, Health promotional messages generated by humans
3 · Its place in the literature

Who cites it

357 citing papers in PubMed, 5 syntheses or guidelines pooled it.

  1. Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology · 2026
    Pooled it
  2. Pooled it
  3. Pooled it
  4. Pooled it
  5. Pooled it
  6. Trial
  7. Trial
  8. Trial
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Validating LLM judges for automated oversight of patient communication.medRxiv : the preprint server for health sciences · 2026
    Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article

297 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

19 authors.

Suhana BediDepartment of Biomedical Data Science, Stanford School of Medicine, Stanford, California.
Yutong LiuClinical Excellence Research Center, Stanford University, Stanford, California.
Lucy Orr-EwingClinical Excellence Research Center, Stanford University, Stanford, California.
Dev DashClinical Excellence Research Center, Stanford University, Stanford, California.
Sanmi KoyejoDepartment of Computer Science, Stanford University, Stanford, California.
Alison CallahanCenter for Biomedical Informatics Research, Stanford University, Stanford, California.
Jason A FriesCenter for Biomedical Informatics Research, Stanford University, Stanford, California.
Michael WornowCenter for Biomedical Informatics Research, Stanford University, Stanford, California.
Akshay SwaminathanCenter for Biomedical Informatics Research, Stanford University, Stanford, California.
Lisa Soleymani LehmannDepartment of Medicine, Harvard Medical School, Boston, Massachusetts.
Hyo Jung HongDepartment of Anesthesiology, Stanford University, Stanford, California.
Mehr KashyapStanford University School of Medicine, Stanford, California.
Akash R ChaurasiaCenter for Biomedical Informatics Research, Stanford University, Stanford, California.
Nirav R ShahClinical Excellence Research Center, Stanford University, Stanford, California.
Karandeep SinghDigital Health Innovation, University of California San Diego Health, San Diego.
Troy TazbazDigital Health Center of Excellence, US Food and Drug Administration, Washington, DC.
Arnold MilsteinClinical Excellence Research Center, Stanford University, Stanford, California.
Michael A PfefferDepartment of Medicine, Stanford University School of Medicine, Stanford, California.
Nigam H ShahClinical Excellence Research Center, Stanford University, Stanford, California.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: Large language models (LLMs) can assist in various health care activities, but current evaluation approaches may not adequately identify the most useful application areas. Objective: To summarize existing evaluations of LLMs in health care in terms of 5 components: (1) evaluation data type, (2) health care task, (3) natural language processing (NLP) and natural language understanding (NLU) tasks, (4) dimension of evaluation, and (5) medical specialty. Data Sources: A systematic search of PubMed and Web of Science was performed for studies published between January 1, 2022, and February 19, 2024. Study Selection: Studies evaluating 1 or more LLMs in health care. Data Extraction and Synthesis: Three independent reviewers categorized studies via keyword searches based on the data used, the health care tasks, the NLP and NLU tasks, the dimensions of evaluation, and the medical specialty. Results: Of 519 studies reviewed, published between January 1, 2022, and February 19, 2024, only 5% used real patient care data for LLM evaluation. The most common health care tasks were assessing medical knowledge such as answering medical licensing examination questions (44.5%) and making diagnoses (19.5%). Administrative tasks such as assigning billing codes (0.2%) and writing prescriptions (0.2%) were less studied. For NLP and NLU tasks, most studies focused on question answering (84.2%), while tasks such as summarization (8.9%) and conversational dialogue (3.3%) were infrequent. Almost all studies (95.4%) used accuracy as the primary dimension of evaluation; fairness, bias, and toxicity (15.8%), deployment considerations (4.6%), and calibration and uncertainty (1.2%) were infrequently measured. Finally, in terms of medical specialty area, most studies were in generic health care applications (25.6%), internal medicine (16.4%), surgery (11.4%), and ophthalmology (6.9%), with nuclear medicine (0.6%), physical medicine (0.4%), and medical genetics (0.2%) being the least represented. Conclusions and Relevance: Existing evaluations of LLMs mostly focus on accuracy of question answering for medical examinations, without consideration of real patient care data. Dimensions such as fairness, bias, and toxicity and deployment considerations received limited attention. Future evaluations should adopt standardized applications and metrics, use clinical data, and broaden focus to include a wider range of tasks and specialties.

Indexed as

Delivery of Health CareLarge Language ModelsNatural Language ProcessingMedicine

Identifiers

PMID39405325
PMCPMC11480901

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.