Evidence map›Paper›PMID 42452955›Full record

ArticleJournal of medical Internet research2026

Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: Comparative Analysis.

Srinivasagam Prabha, Bernardo Gabriele Collaco, Cesar Abraham Gomez-Cabello, Syed Ali Haider, Ariana Genovese, Zhihui Fang, Nadia Wood, Sanjay Bagaria, Cui Tao, Antonio Jorge Forte

Abstract readComparative Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Srinivasagam PrabhaDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0000-0003-2573-8493
Bernardo Gabriele CollacoDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0000-0003-3845-2646
Cesar Abraham Gomez-CabelloDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0009-0008-0603-3192
Syed Ali HaiderDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0009-0007-5621-2861
Ariana GenoveseDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0009-0000-9678-2163
Zhihui FangDivision of Clinical Trials and Biostatistics, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID 0009-0000-4745-4745
Nadia WoodDepartment of Radiology AI IT, Mayo Clinic, Rochester, MN, United States.ORCID 0000-0002-2685-6427
Sanjay BagariaDepartment of Surgery, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID 0000-0001-6677-8964
Cui TaoDepartment of Artificial Intelligence and Informatics, Mayo Clinic in Florida, Jacksonville, FL, United States.ORCID 0000-0002-4267-1924
Antonio Jorge ForteDivision of Plastic Surgery, Mayo Clinic in Florida, 4500 San Pablo Road South, Jacksonville, FL, 32224, United States, 1 904-953-2000.ORCID 0000-0003-2004-7538

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) show growing potential for decision support. However, integrating domain-specific medical knowledge while maintaining accuracy, safety, and interpretability remains challenging for postoperative discharge instructions and patient education. Fine-tuning, retrieval-augmented generation (RAG), and hybrid fine-tuning+RAG approaches are prominent strategies for knowledge integration, but their comparative performance in postoperative care has not been systematically evaluated. Objective: We aimed to compare the performance, reliability, and safety characteristics of baseline, fine-tuning, RAG, and hybrid fine-tuning+RAG LLM configurations for postoperative decision support. Methods: We conducted a comparative evaluation of 4 LLM configurations using Google Gemini 2.5 Flash. A total of 600 postoperative question and answer pairs were used for model adaptation and validation, while 150 queries were reserved for final evaluation. Queries included routine postoperative questions, emergency escalation scenarios, and deliberately out-of-scope prompts. Outputs were independently assessed by 3 blinded clinical experts for clinical medical accuracy, safety or refusal accuracy, completeness, and relevance. Automated metrics evaluated readability, faithfulness, and hallucination propensity. Results: All knowledge-enhanced models significantly outperformed baseline in overall accuracy (baseline 68% vs fine-tuning 92.7%, RAG 91.3%, fine-tuning+RAG 97.3%; P<.001). For in-scope clinical queries, fine-tuning+RAG achieved the highest clinical medical accuracy (96.7%) and was the only configuration to significantly outperform baseline in pairwise comparisons. Enhanced models also demonstrated higher safety or refusal accuracy than baseline; however, the baseline configuration did not receive equivalent safety or deferral instructions, which likely influenced these findings. Fine-tuning+RAG achieved the strongest composite classification performance, including 100% precision, 96.7% recall, and 98.3% F1-score. Fine-tuning and RAG showed broadly comparable performance across most secondary outcomes. Although knowledge-enhanced models demonstrated lower readability than baseline, restricted analysis of 100 routine in-scope postoperative queries suggested that part of this difference was attributable to standardized safety boilerplate. Conclusions: Incorporating domain-specific knowledge through fine-tuning, RAG, or both improved postoperative decision-support performance compared with the baseline LLM. All knowledge-enhanced approaches demonstrated strong performance, with the hybrid fine-tuning+RAG configuration achieving the most favorable overall point estimates across several outcomes. However, differences among the enhanced configurations were generally modest and less evident in sensitivity analyses restricted to unanimously rated queries. These findings support knowledge-enhanced LLMs as promising tools for postoperative education and decision support, while highlighting the need for further validation, readability optimization, transparent governance, and sustained human oversight before patient-facing deployment.

Indexed as

Decision Support Systems, ClinicalPostoperative CareHumansLarge Language ModelsReproducibility of ResultsAIdecision supportfine-tuninglarge language modelspostoperative careretrieval-augmented generation

Identifiers

PMID42452955
PMCPMC13369304

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.