Evidence map›Paper›PMID 40615349›Full record

ArticleThe Journal of international medical research2025

Clinical applications of large language models in medicine and surgery: A scoping review.

Eric Nan Liang, Sophia Pei, Phillip Staibano, Benjamin van der Woerd

Abstract readScoping Review
In one paragraph

Article in The Journal of international medical research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Article
  3. Artificial Intelligence in Rhinology: A State-of-the-Art Review of Clinical Readiness and Implementation Pathways.Otolaryngology--head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery · 2026
    Review
  4. LLM-powered prostate cancer staging from PSMA-PET/CT reports using PROMISE v2.European journal of nuclear medicine and molecular imaging · 2026
    Article
  5. Review
  6. Article
  7. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Eric Nan LiangMichael G. DeGroote School of Medicine, McMaster University, Canada.ORCID 0009-0005-6702-3957
Sophia PeiMichael G. DeGroote School of Medicine, McMaster University, Canada.
Phillip StaibanoDivision of Otolaryngology-Head and Neck Surgery, Department of Surgery, McMaster University, Canada.
Benjamin van der WoerdDivision of Otolaryngology-Head and Neck Surgery, Department of Surgery, McMaster University, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

ObjectiveTo provide a comprehensive overview of the current use of large language models in clinical medicine and surgery, with emphasis on model characteristics, clinical applications, and readiness for adoption.MethodsA scoping review of studies on the use of large language models in clinical medicine and surgery was conducted in accordance with the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA)-scoping review and JBI methodology (protocol registration: 10.37766/inplasy2025.3.0102). A comprehensive search of EMBASE, PubMed, CINAHL, and IEEE Xplore identified 3313 articles published between 2018 and 2023. After screening of articles and full-text review, 156 studies were included. Data were extracted for study type, sample size, clinical specialty, model architecture, training methods, application purpose, and performance metrics. Descriptive analyses were performed.ResultsMost studies were proof-of-concept studies (55.8%) or clinical trials (21.2%), with a steady rise in publications since 2022. Large language models were most frequently used for data extraction (69.9%), followed by clinical recommendations (11.5%), report generation (9.0%), and patient-facing chatbots (7.1%). Proprietary models were used in 57.7% of the studies, whereas 39.7% used open-source models. ChatGPT-3.5, ChatGPT-4, and Bidirectional Encoder Representations from Transformers (BERT) were the most commonly reported models. Only 25.0% of the studies reported models as ready for clinical use, whereas 67.9% stated that the models required further validation. F-score (30.8%) and area under the curve (15.4%) were the most common performance metrics; 10.9% of the studies used expert opinion for validation.ConclusionsLarge language models are increasingly being used in clinical medicine. Although most applications focus on data extraction and summarization, emerging studies are beginning to explore higher-level tasks such as clinical decision-making and multidisciplinary simulation. Significant heterogeneity continues to exist in model architecture, evaluation methods, and reporting standards. Further standardization is needed to develop transparent evaluation frameworks and ensure safe, reliable integration of large language models into complex clinical workflows.

Indexed as

Clinical MedicineLanguageHumansLarge Language Modelsartificial intelligencegenerative pre-trained transformerLarge language modelsmachine learningnatural language processing

Identifiers

PMID40615349
PMCPMC12227933

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.