SynthesisJournal of the American Medical Informatics Association : JAMIA2025
Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines.
Synthesis in Journal of the American Medical Informatics Association : JAMIA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 76 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
76 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Generative large language models in the clinical management of Alzheimer's disease and mild cognitive impairment.Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical Neurophysiology · 2026Pooled it
- Improving reproducibility of Plasmodium falciparum culture: a large-language-model-driven literature review and a parasite growth variability assessment of donor red blood cells.Malaria journal · 2025Pooled it
- Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A Blinded Expert-Rated Comparative Study of 11 Systems.Journal of clinical periodontology · 2026Article
- Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports using Large Language Models.medRxiv : the preprint server for health sciences · 2026Article
- Large Language Models in Preclinical Spine Research: A Scoping Review and Expert Perspective on Evidence-Aware Experimental Workflows.JOR spine · 2026Article
- Large language models for pneumonia detection in radiology reports via text analysis.PLOS digital health · 2026Article
- SynBioGPT2: A dynamic reasoning framework enables high-fidelity design of microbial cell factories.Biodesign research · 2026Article
- The economics of accuracy for medical reasoning with large language models.PLOS digital health · 2026Article
- Ensuring trustworthy AI assisted guideline development for clinical practice.NPJ digital medicine · 2026Article
- Article
- Large Language Model-Based Clinical Decision Support for Antibiotic Selection and Dose Recommendation in Hospitalized Patients With Pneumonia: Multicenter Retrospective Study.JMIR medical informatics · 2026Article
- A Self-Controlled Benchmark of Retrieval-Augmented Generation for Large Language Models on Clinical Guideline Questions.Diagnostics (Basel, Switzerland) · 2026Article
- Embracing Large Language Models for Medical Applications, Part II: Building a Framework for Clinical Stewardship.Cureus · 2026Article
- Automated Review of Patient Records: Privacy-Preserving Large Language Models for Identifying Incident Nonarteritic Anterior Ischemic Optic Neuropathy at Scale.Ophthalmology science · 2026Article
- Large reasoning models as thinking machines for medicine.Nature biomedical engineering · 2026Review
- Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: Comparative Analysis.Journal of medical Internet research · 2026Article
- Leveraging generative artificial intelligence for the development of non-interventional research study protocols: a proof-of-concept feasibility study.BMC medical research methodology · 2026Article
- Artificial stupidity or logimorphism? How misuse of language warps our thinking about 'artificial intelligence'.European heart journal. Digital health · 2026Article
- Methods to build and maintain effective clinician-patient relationships for adults with disorders of gut-brain interaction (DGBI): a systematic review.Journal of gastroenterology · 2026Review
- Retrieval-Augmented Language Models for Clinical Decision Support in the Classification of Inborn Errors of Immunity.Journal of clinical immunology · 2026Article
16 more citing papers are in PubMed but not listed here.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
Abstract
objectiveThe objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness. MATERIALS AND
methodsWe conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to "retrieval augmented generation" and "large language model," for articles published in 2023 and 2024. We selected studies that compared baseline LLM performance with RAG performance. We developed a random-effect meta-analysis model, using odds ratio as the effect size.
resultsAmong 335 studies, 20 were included in this literature review. The pooled effect size was 1.35, with a 95% confidence interval of 1.19-1.53, indicating a statistically significant effect (P = .001). We reported clinical tasks, baseline LLMs, retrieval sources and strategies, as well as evaluation methods. DISCUSSION: Building on our literature review, we developed Guidelines for Unified Implementation and Development of Enhanced LLM Applications with RAG in Clinical Settings to inform clinical applications using RAG.
conclusionOverall, RAG implementation showed a 1.35 odds ratio increase in performance compared to baseline LLMs. Future research should focus on (1) system-level enhancement: the combination of RAG and agent, (2) knowledge-level enhancement: deep integration of knowledge into LLM, and (3) integration-level enhancement: integrating RAG systems within electronic health records.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.