ArticleJournal of medical Internet research2026
Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Large Language Models for Postoperative Decision Support: Comparative Analysis.
Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Large language models (LLMs) show growing potential for decision support. However, integrating domain-specific medical knowledge while maintaining accuracy, safety, and interpretability remains challenging for postoperative discharge instructions and patient education. Fine-tuning, retrieval-augmented generation (RAG), and hybrid fine-tuning+RAG approaches are prominent strategies for knowledge integration, but their comparative performance in postoperative care has not been systematically evaluated. Objective: We aimed to compare the performance, reliability, and safety characteristics of baseline, fine-tuning, RAG, and hybrid fine-tuning+RAG LLM configurations for postoperative decision support. Methods: We conducted a comparative evaluation of 4 LLM configurations using Google Gemini 2.5 Flash. A total of 600 postoperative question and answer pairs were used for model adaptation and validation, while 150 queries were reserved for final evaluation. Queries included routine postoperative questions, emergency escalation scenarios, and deliberately out-of-scope prompts. Outputs were independently assessed by 3 blinded clinical experts for clinical medical accuracy, safety or refusal accuracy, completeness, and relevance. Automated metrics evaluated readability, faithfulness, and hallucination propensity. Results: All knowledge-enhanced models significantly outperformed baseline in overall accuracy (baseline 68% vs fine-tuning 92.7%, RAG 91.3%, fine-tuning+RAG 97.3%; P<.001). For in-scope clinical queries, fine-tuning+RAG achieved the highest clinical medical accuracy (96.7%) and was the only configuration to significantly outperform baseline in pairwise comparisons. Enhanced models also demonstrated higher safety or refusal accuracy than baseline; however, the baseline configuration did not receive equivalent safety or deferral instructions, which likely influenced these findings. Fine-tuning+RAG achieved the strongest composite classification performance, including 100% precision, 96.7% recall, and 98.3% F1-score. Fine-tuning and RAG showed broadly comparable performance across most secondary outcomes. Although knowledge-enhanced models demonstrated lower readability than baseline, restricted analysis of 100 routine in-scope postoperative queries suggested that part of this difference was attributable to standardized safety boilerplate. Conclusions: Incorporating domain-specific knowledge through fine-tuning, RAG, or both improved postoperative decision-support performance compared with the baseline LLM. All knowledge-enhanced approaches demonstrated strong performance, with the hybrid fine-tuning+RAG configuration achieving the most favorable overall point estimates across several outcomes. However, differences among the enhanced configurations were generally modest and less evident in sensitivity analyses restricted to unanimously rated queries. These findings support knowledge-enhanced LLMs as promising tools for postoperative education and decision support, while highlighting the need for further validation, readability optimization, transparent governance, and sustained human oversight before patient-facing deployment.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.