Evidence map›Paper›PMID 39812777›Full record

SynthesisJournal of the American Medical Informatics Association : JAMIA2025

Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines.

Siru Liu, Allison B McCoy, Adam Wright

Abstract readSystematic ReviewMeta-Analysis
In one paragraph

Synthesis in Journal of the American Medical Informatics Association : JAMIA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 76 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
76citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

76 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology · 2026
    Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Review
  16. Article
  17. Article
  18. Article
  19. Review
  20. Article

16 more citing papers are in PubMed but not listed here.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Siru LiuDepartment of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37212, United States.
Allison B McCoyDepartment of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37212, United States.ORCID 0000-0003-2292-9147
Adam WrightDepartment of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37212, United States.ORCID 0000-0001-6844-145X

Funding

Strategies for Engineering Reliable Value Sets (SERVS)R01LM013995 · NLM · VANDERBILT UNIVERSITY MEDICAL CENTER · PI ADAM T WRIGHT · 2022 to 2026
$1.0M
Optimizing Clinical Decision Support Alerts Using Explainable Artificial Intelligence (XAI)R00LM014097 · NLM · VANDERBILT UNIVERSITY MEDICAL CENTER · PI LIU, SIRU · 2023 to 2024
$498k
NIH HHS R00LM014097-02NLM NIH HHS R00 LM014097NLM NIH HHS R01 LM013995
6 · The paper itself

Abstract

objectiveThe objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness. MATERIALS AND

methodsWe conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to "retrieval augmented generation" and "large language model," for articles published in 2023 and 2024. We selected studies that compared baseline LLM performance with RAG performance. We developed a random-effect meta-analysis model, using odds ratio as the effect size.

resultsAmong 335 studies, 20 were included in this literature review. The pooled effect size was 1.35, with a 95% confidence interval of 1.19-1.53, indicating a statistically significant effect (P = .001). We reported clinical tasks, baseline LLMs, retrieval sources and strategies, as well as evaluation methods. DISCUSSION: Building on our literature review, we developed Guidelines for Unified Implementation and Development of Enhanced LLM Applications with RAG in Clinical Settings to inform clinical applications using RAG.

conclusionOverall, RAG implementation showed a 1.35 odds ratio increase in performance compared to baseline LLMs. Future research should focus on (1) system-level enhancement: the combination of RAG and agent, (2) knowledge-level enhancement: deep integration of knowledge into LLM, and (3) integration-level enhancement: integrating RAG systems within electronic health records.

Indexed as

Information Storage and RetrievalNatural Language ProcessingProgramming LanguagesHumansLarge Language Modelslarge language modelmeta-analysisretrieval augmented generationsystematic review

Identifiers

PMID39812777
PMCPMC12005634

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.