SynthesisInternational journal of medical informatics2026
Performance and improvement strategies for adapting generative large language models for electronic health record applications: A systematic review.
Synthesis in International journal of medical informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
10 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Testing and evaluation of generative large language models in electronic health record applications: a systematic review.Journal of the American Medical Informatics Association : JAMIA · 2026Pooled it
- Generative large language models in medicine: a scoping review of recent methodological advances.npj health systems · 2026Review
- The Performance of Large Language Models in Extracting Intestinal Symptoms From Electronic Health Records: Retrospective Observational Study.Journal of medical Internet research · 2026Observational
- Precision pharmacology: deep learning infused ontological framework with E-GRU enhancement for tailored medicine prescriptions.Scientific reports · 2026Article
- An automatic consult reply system for therapeutic plasma exchange using retrieval-augmented generation.Vox sanguinis · 2026Article
- [Large language models as a communication and organizational infrastructure in urology: evidence, limitations, and clinical responsibility].Urologie (Heidelberg, Germany) · 2026Review
- irAE-GPT: leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets.EBioMedicine · 2026Article
- Structured taxonomy and framework for developing medical benchmark in large language models derived from scoping review.NPJ digital medicine · 2026Article
- Evaluating Chain-of-Thought reasoning in large language models for thyroid ultrasound interpretation: a dual-information approach.Frontiers in artificial intelligence · 2026Article
- TCMEval-PA: a question-answering benchmark dataset for the prescription audit of Traditional Chinese Medicine.Scientific data · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
Abstract
purposeTo synthesize performance and improvement strategies for adapting generative LLMs in EHR analyses and applications.
methodsWe followed the PRISMA guidelines to conduct a systematic review of articles from PubMed and Web of Science published between January 1, 2023 and November 9, 2024. Multiple reviewers including biomedical informaticians and a clinician involved in the article reviewing process. Studies were included if they used generative LLMs to analyze real-world EHR data and reported quantitative performance evaluations for an improvement technique. The review identified key clinical applications, summarized performance and the improvement strategies.
resultsOf the 18,735 articles retrieved, 196 met our criteria. 112 (57.1%) studies used generative LLMs for clinical decision support tasks, 40 (20.4%) studies involved documentation tasks, 39 (19.9%) studies involved information extraction tasks, 11 (5.6%) studies involved patient communication tasks, and 10 (5.1%) studies included summarization tasks. Among the 196 studies, most studies (88.8%) did not quantitatively evaluate the LLM performance improvement strategies, with the rest twenty-four studies (12.2%) quantitatively evaluated the effectiveness of in-context learning (9 studies), fine-tuning (12 studies), multimodal integration (8 studies), and ensemble learning (2 studies). Three studies highlighted that few-shot prompting, fine-tuning, and multimodal data integration might not improve performance, and another two studies found that fine-tuning a smaller model could outperform a large model.
conclusionApplying a performance improvement strategy may not necessarily lead to performance improvement, and detailed guidelines regarding how to apply those strategies more effectively and safely are needed, which can be completed from more quantitative analysis in the future.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.