Evidence map›Paper›PMID 40885071›Full record

SynthesisInternational journal of medical informatics2026

Performance and improvement strategies for adapting generative large language models for electronic health record applications: A systematic review.

Xinsong Du, Zhengyang Zhou, Yifei Wang, Ya-Wen Chuang, Yiming Li, Richard Yang, Pengyu Hong, David W Bates, Li Zhou

Abstract readSystematic Review
In one paragraph

Synthesis in International journal of medical informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Review
  3. Observational
  4. Article
  5. Article
  6. Review
  7. Article
  8. Article
  9. Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Xinsong DuDivision of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States; Department of Medicine, Harvard Medical School, Boston, MA 02115, United States. Electronic address: xidu1@bwh.harvard.edu.
Zhengyang ZhouDepartment of Computer Science, Brandeis University, Waltham, MA 02453, United States.
Yifei WangDepartment of Computer Science, Brandeis University, Waltham, MA 02453, United States.
Ya-Wen ChuangDivision of Nephrology, Department of Internal Medicine, Taichung Veterans General Hospital, Taichung 407219, Taiwan; Department of Post-Baccalaureate Medicine, College of Medicine, National Chung Hsing University, Taichung 402202, Taiwan; School of Medicine, College of Medicine, China Medical University, Taichung 404328, Taiwan.
Yiming LiDivision of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States; Department of Medicine, Harvard Medical School, Boston, MA 02115, United States.
Richard YangDivision of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States; Department of Medicine, Harvard Medical School, Boston, MA 02115, United States.
Pengyu HongDepartment of Computer Science, Brandeis University, Waltham, MA 02453, United States.
David W BatesDivision of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States; Department of Medicine, Harvard Medical School, Boston, MA 02115, United States; Department of Health Policy and Management, Harvard T.H. Chan School of Public Health, Boston, MA 02115, United States.
Li ZhouDivision of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA 02115, United States; Department of Medicine, Harvard Medical School, Boston, MA 02115, United States.

Funding

Leveraging Longitudinal Data and Informatics Technology to Understand the Role of Bilingualism in Cognitive Resilience, Aging and DementiaR01AG080429 · NIA · BRIGHAM AND WOMEN'S HOSPITAL · PI Michelle L Dossett, HUA XU · 2023 to 2026
$5.5M
Identifying and addressing missingness and bias to enhance discovery from multimodal health dataR01LM014239 · NLM · BRIGHAM AND WOMEN'S HOSPITAL · PI Pengyu Hong, Li Zhou · 2023 to 2026
$1.6M
NIA NIH HHS R01 AG080429NLM NIH HHS R01 LM014239
6 · The paper itself

Abstract

purposeTo synthesize performance and improvement strategies for adapting generative LLMs in EHR analyses and applications.

methodsWe followed the PRISMA guidelines to conduct a systematic review of articles from PubMed and Web of Science published between January 1, 2023 and November 9, 2024. Multiple reviewers including biomedical informaticians and a clinician involved in the article reviewing process. Studies were included if they used generative LLMs to analyze real-world EHR data and reported quantitative performance evaluations for an improvement technique. The review identified key clinical applications, summarized performance and the improvement strategies.

resultsOf the 18,735 articles retrieved, 196 met our criteria. 112 (57.1%) studies used generative LLMs for clinical decision support tasks, 40 (20.4%) studies involved documentation tasks, 39 (19.9%) studies involved information extraction tasks, 11 (5.6%) studies involved patient communication tasks, and 10 (5.1%) studies included summarization tasks. Among the 196 studies, most studies (88.8%) did not quantitatively evaluate the LLM performance improvement strategies, with the rest twenty-four studies (12.2%) quantitatively evaluated the effectiveness of in-context learning (9 studies), fine-tuning (12 studies), multimodal integration (8 studies), and ensemble learning (2 studies). Three studies highlighted that few-shot prompting, fine-tuning, and multimodal data integration might not improve performance, and another two studies found that fine-tuning a smaller model could outperform a large model.

conclusionApplying a performance improvement strategy may not necessarily lead to performance improvement, and detailed guidelines regarding how to apply those strategies more effectively and safely are needed, which can be completed from more quantitative analysis in the future.

Indexed as

Electronic Health RecordsNatural Language ProcessingDecision Support Systems, ClinicalHumansLarge Language ModelsAIChatGPTEHREvaluationLLMPerformanceReview

Identifiers

PMID40885071
PMCPMC12413914

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.