ReviewMayo Clinic proceedings. Digital health2025
Fine-Tuning Large Language Models for Specialized Use Cases.
Review in Mayo Clinic proceedings. Digital health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 37 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
37 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review.Journal of medical Internet research · 2026Pooled it
- Large language models in real-world clinical workflows: a systematic review of applications and implementation.Frontiers in digital health · 2025Pooled it
- Article
- Review
- Sequential multi-site fine-tuning for incremental deployment of large language models for mobility functional status extraction.JAMIA open · 2026Article
- Tutorial: guidance on the use of large language models for medical research.Nature protocols · 2026Review
- Supervised Fine-Tuning of Large Language Models With Chain-of-Thought Reasoning for Pediatric Heart Disease Detection in Unstructured Echocardiogram Reports: Algorithm Development and Validation.JMIR formative research · 2026Article
- Large Language Models and Their Applications in Mental Health: Scoping Review.JMIR mental health · 2026Article
- Assessing the accuracy and educational value of ChatGPT-generated content for core topics in cardiology: a descriptive analysis at Selçuk University Cardiology Clinic.BMC medical education · 2026Article
- Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability.JCO clinical cancer informatics · 2026Article
- Domain-adapted language model using reinforcement learning for various dementias.medRxiv : the preprint server for health sciences · 2026Article
- AF-CuRL: Stable Reinforcement Learning for Resource-Constrained Long-Form Reasoning in Edge-Intelligent Systems.Sensors (Basel, Switzerland) · 2026Article
- ChatGPT models provide higher-quality but lower-readability responses than Google Gemini regarding anterior shoulder instability, with no added benefit of the orthopaedic expert plugin.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- Scaling medical AI across clinical contexts.Nature medicine · 2026Review
- Assessing multiple-choice question quality in internal medicine: a comparative analysis of three large language models against expert consensus.Frontiers in medicine · 2026Article
- Extracting language information from clinical notes using large language models.International journal of medical informatics · 2026Article
- Performance and safety of a fine-tuned small language model for pediatric emergency triage: A benchmark study.PloS one · 2026Article
- Token-splitting improves GPT-4.1 performance on plastic surgery exams: implications for AI-Assisted medical education.Medical education online · 2025Article
- Does DeepSeek Provide Clinically Acceptable Intraocular Lens (IOL) Power Predictions in Cataract Surgery? A Proof-of-Concept Study.Journal of clinical medicine · 2025Article
- Can general purpose large language models assist pediatricians in predicting infants with serious bacterial infection?BMC medical informatics and decision making · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Large language models (LLMs) are a type of artificial intelligence, which operate by predicting and assembling sequences of words that are statistically likely to follow from a given text input. With this basic ability, LLMs are able to answer complex questions and follow extremely complex instructions. Products created using LLMs such as ChatGPT by OpenAI and Claude by Anthropic have created a huge amount of traction and user engagements and revolutionized the way we interact with technology, bringing a new dimension to human-computer interaction. Fine-tuning is a process in which a pretrained model, such as an LLM, is further trained on a custom data set to adapt it for specialized tasks or domains. In this review, we outline some of the major methodologic approaches and techniques that can be used to fine-tune LLMs for specialized use cases and enumerate the general steps required for carrying out LLM fine-tuning. We then illustrate a few of these methodologic approaches by describing several specific use cases of fine-tuning LLMs across medical subspecialties. Finally, we close with a consideration of some of the benefits and limitations associated with fine-tuning LLMs for specialized use cases, with an emphasis on specific concerns in the field of medicine.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.