Evidence map›Paper›PMID 42101844›Full record

ArticleJAMA network open2026

Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries.

François Grolleau, April S Liang, Timothy Keyes, Stephen P Ma, Thomas Lew, Tridu R Huynh, Natasha Steele, Philip Chung, Paige Qin, Gowri Chandra and 13 more

Abstract read
In one paragraph

Article in JAMA network open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

23 authors.

François GrolleauDivision of Computational Medicine, Stanford University, Stanford, California.
April S LiangDivision of Hospital Medicine, Stanford University, Stanford, California.
Timothy KeyesTechnology and Digital Solutions, Stanford Health Care, Palo Alto, California.
Stephen P MaDivision of Hospital Medicine, Stanford University, Stanford, California.
Thomas LewDivision of Hospital Medicine, Stanford University, Stanford, California.
Tridu R HuynhDivision of Hospital Medicine, Stanford University, Stanford, California.
Natasha SteeleDivision of Hospital Medicine, Stanford University, Stanford, California.
Philip ChungDepartment of Anesthesiology, Perioperative, and Pain Medicine, Stanford University, Stanford, California.
Paige QinDivision of Hospital Medicine, Stanford University, Stanford, California.
Gowri ChandraDivision of Hospital Medicine, Stanford University, Stanford, California.
Stephanie F WangDivision of Hospital Medicine, Stanford University, Stanford, California.
Evan MullenDivision of Hospital Medicine, Stanford University, Stanford, California.
Lauren CarpenterDivision of Hospital Medicine, Stanford University, Stanford, California.
Mita HoppenfeldDivision of Hospital Medicine, Stanford University, Stanford, California.
Matthew MorrinDivision of Hospital Medicine, Stanford University, Stanford, California.
Baffour A KyerematenDivision of Hospital Medicine, Stanford University, Stanford, California.
Nerissa AmbersTechnology and Digital Solutions, Stanford Health Care, Palo Alto, California.
Nikesh KotechaTechnology and Digital Solutions, Stanford Health Care, Palo Alto, California.
Emily AlsentzerDepartment of Biomedical Data Science, Stanford University, Stanford, California.
Jason HomDivision of Hospital Medicine, Stanford University, Stanford, California.
Nigam H ShahDivision of Computational Medicine, Stanford University, Stanford, California.
Kevin SchulmanDivision of Hospital Medicine, Stanford University, Stanford, California.
Jonathan H ChenDivision of Computational Medicine, Stanford University, Stanford, California.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest that large language models (LLMs) can generate clinical summaries of comparable quality to those by physicians, prospective data on their safety, utility, and association with clinician well-being in clinical environments are lacking. Objective: To evaluate the safety, use, and association with clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries, during prospective clinical deployment. Design, Setting, and Participants: This single-arm prospective pilot quality improvement study encompassed hospital discharges at 1 academic inpatient medicine unit from August 1 to October 11, 2025, with baseline comparisons drawn from April 9 to July 31, 2025. Intervention: A custom agentic LLM workflow using Gemini 2.5 Pro generated draft hospital course summaries nightly using patient history and physical and daily progress notes. Drafts were securely emailed to physicians daily for review and optional use. Main Outcomes and Measures: The primary outcome was physician-reported potential for and severity of harm from unedited summaries (Agency for Healthcare Research and Quality Common Format Harm Scale). Secondary outcomes included use rate, error types (omissions, inaccuracies, and hallucinations), time spent in discharge summaries (electronic health record logs), and changes in cognitive burden (NASA Task Load Index; score range, 0-100, with higher scores indicating greater cognitive burden) and burnout (Stanford Professional Fulfillment Index Work Exhaustion Scale; score range, 0-4, with higher scores indicating greater burnout). Results: Among 384 hospital discharges, the system generated 1274 summaries. Physicians used artificial intelligence (AI) content in 219 cases (57.0%). Feedback on 100 summaries (88 of 219 used summaries [40.2%] and 12 of 165 unused summaries [7.3%]) noted omissions (25 summaries [25.0%]) and inaccuracies (20 summaries [20.0%]) but rare hallucinations (2 summaries [2.0%]). Physicians rated 88 unedited summaries (88.0%) as having no harm potential and 1 (1.0%) as likely to cause moderate harm; no severe harm was reported. Mean physician burnout scores decreased significantly from before to after the intervention (1.75; 95% CI, 1.16-2.34 vs 1.20; 95% CI, 0.71-1.69; P = .03). Time savings were heterogeneous, with 5 of 7 physicians with matched baseline data (71.4%) seeing reductions in median documentation time; changes from baseline to pilot were up to 2.9 minutes, which was a nonsignificant difference (10.7 minutes; 95% CI, 7.4-13.3 minutes vs 7.8 minutes; 95% CI, 5.1-11.7 minutes; P = .13). Conclusions and Relevance: In this study, an LLM-based agentic workflow produced hospital course summaries that were frequently used with minimal risk of harm identified. The intervention was associated with a reduction in physician burnout, supporting the viability of AI summarization to mitigate documentation burden.

Indexed as

Patient Discharge SummariesPatient SafetyPhysiciansElectronic Health RecordsGenerative Artificial IntelligenceHumansLarge Language ModelsPilot ProjectsProspective StudiesQuality Improvement

Identifiers

PMID42101844
PMCPMC13156793

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.