Evidence map›Paper›PMID 42753254›Full record

SynthesisJournal of medical Internet research2026

Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review.

Anshum Patel, Yugant Khand, Sai Krishna Vallamchetla, Pengze Li, Cui Tao, Joseph Cheung

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Anshum PatelDivision of Pulmonary, Allergy and Sleep Medicine, Mayo Clinic, 4500 San Pablo Road South, Jacksonville, FL, United States, 1 904-953-2000.
Yugant KhandDepartment of Neurology, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID http://orcid.org/0000-0002-3057-3200
Sai Krishna VallamchetlaDepartment of Neurology, Mayo Clinic in Florida, Jacksonville, FL, United States.
Pengze LiDepartment of Artificial Intelligence and Informatics, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID http://orcid.org/0000-0001-7015-0491
Cui TaoDepartment of Artificial Intelligence and Informatics, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID http://orcid.org/0000-0002-4267-1924
Joseph CheungDivision of Pulmonary, Allergy and Sleep Medicine, Mayo Clinic, 4500 San Pablo Road South, Jacksonville, FL, United States, 1 904-953-2000.ORCID http://orcid.org/0000-0002-1222-7603

Funding

ACTS (AD Clinical Trial Simulation): Developing Advanced Informatics Approaches for an Alzheimer's Disease Clinical Trial Simulation SystemR01AG084236 · NIA · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI Jiang Bian, Cui Tao · 2023 to 2026
$4.1M
Standardizing and Harmonizing Behavioral and Social Science Research Factors in Alzheimer's Disease through Ontology-Based ApproachesU01AG088076 · NIA · MAYO CLINIC JACKSONVILLE · PI Jiang Bian, Cui Tao · 2024 to 2026
$2.3M
NIA NIH HHS R01 AG084236NIA NIH HHS U01 AG088076
6 · The paper itself

Abstract

Background: Large language models (LLMs) demonstrate strong performance on medical knowledge benchmarks, but their safe and effective use in clinical practice depends on posttraining adaptation rather than raw model capability. Fine-tuning, retrieval-augmented generation (RAG), and hybrid approaches are principal strategies for grounding language models in clinical evidence, yet their comparative effectiveness remains unclear. Objective: This systematic review aims to synthesize evidence on fine-tuning, RAG, and hybrid posttraining strategies for clinical diagnosis and decision-support tasks and to identify strategy-task alignments and methodological features associated with improved performance. Methods: We conducted a systematic review in accordance with PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. PubMed/MEDLINE, Scopus, and Web of Science were searched from January 2018 through May 2026. Eligible studies evaluated transformer-based language models that underwent posttraining adaptation, retrieval augmentation, or both for clinical decision support, diagnosis, triage, risk stratification, or related health care applications. Studies evaluating nonadapted models, non-language-model AI systems, prompt engineering without performance evaluation, or nonclinical applications were excluded. Data extracted included model architecture, adaptation strategy, clinical domain, validation approach, and performance outcomes. Risk of bias was assessed using PROBAST+AI (Prediction model Risk of Bias Assessment Tool for AI). Studies were grouped according to the primary enhancement strategy (fine-tuning or parameter-efficient fine-tuning, RAG, or hybrid approaches), and findings were synthesized descriptively. Results: Of 1890 identified records, 35 studies published between 2024 and 2026 met eligibility criteria. Enhancement strategies included RAG (17/35, 48.6%), fine-tuning or parameter-efficient fine-tuning (7/35, 20%), and hybrid approaches (11/35, 31.4%). Studies included diverse specialties from oncology, neurology, radiology, mental health, cardiology, ophthalmology, and surgical care. Fine-tuning demonstrated strong performance for task-specific applications, achieving area under the receiver operating characteristic curve values up to 0.912 for cancer detection and area under curve of 0.892 for major depressive disorder prediction, while matching clinician-level diagnostic performance in several studies. RAG improved guideline adherence and diagnostic accuracy, with increases from 71.1% to 92.1% and from 78.9% to 94.7% in guideline-based decision-support tasks. However, benefits were inconsistent across larger reasoning-capable models. Hybrid systems generally achieved the strongest performance in complex clinical workflows, with external validation accuracies exceeding 90% in stroke triage, dermatology, multimodal imaging, and oncology applications. Risk-of-bias assessment identified substantial methodological limitations, with 25 studies judged as high risk, 9 as unclear risk, and only 1 as low risk overall. Common concerns included inadequate external validation, lack of calibration assessment, nonrepresentative participant selection, and insufficient reporting of analytical methods. Conclusions: Adaptation strategies should align with task needs, using fine-tuning for narrow classification, RAG for guideline-grounded reasoning, and hybrid approaches for complex multimodal tasks. However, the evidence base remains largely retrospective or benchmark-based. Prospective studies with external validation, calibration, and standardized safety reporting are needed before broader clinical use.

Indexed as

Clinical Decision-MakingDecision Support Systems, ClinicalDelivery of Health CareLarge Language ModelsHumansAIAI agentsclinical decision supportevidence-based medicinefine-tuninggenerative AIhealth carelanguage modelslarge language modelsLLMsnatural language processingposttraining methodsRAGreinforcement learningretrieval-augmented generationSLMssmall language models

Identifiers

PMID42753254
PMCPMC13585368

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.