Evidence map›Paper›PMID 40597031›Full record

ArticleBMC medical education2025

Large language models versus traditional textbooks: optimizing learning for plastic surgery case preparation.

Chandler Hinson, Cybil Sierra Stingl, Rahim Nazerali

Abstract readComparative Study
In one paragraph

Article in BMC medical education, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Chandler HinsonFrederick P. Whiddon College of Medicine, University of South Alabama, 5851 USA North Drive, Mobile, AL, 36688, USA. csh2121@jagmail.southalabama.edu.
Cybil Sierra StinglStanford Department of Surgery, Division of Plastic and Reconstructive Surgery, 770 Welch Road, Palo Alto, CA, 94304, USA.
Rahim NazeraliStanford Department of Surgery, Division of Plastic and Reconstructive Surgery, 770 Welch Road, Palo Alto, CA, 94304, USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs), such as ChatGPT-4 and Gemini, represent a new frontier in surgical education by offering dynamic, interactive learning experiences. Despite their potential, concerns about the accuracy, depth of knowledge, and bias in LLM responses persist. This study evaluates the effectiveness of LLMs in aiding surgical trainees in plastic and reconstructive surgery through comparison with traditional case-preparation textbooks.

methodsSix representative cases from key areas of plastic and reconstructive surgery-craniofacial, hand, microsurgery, burn, gender-affirming, and aesthetics-were selected. Four types of questions were developed for each case to cover clinical anatomy, indications, contraindications, and complications. Responses from LLMs (ChatGPT-4 and Gemini) and textbooks were compared using surveys distributed to medical students, research fellows, residents, and attending surgeons. Reviewers rated each response on accuracy, thoroughness, usefulness for case preparation, brevity, and overall quality using a 5-point Likert scale. Statistical analyses, including ANOVA and unpaired T-tests, were conducted to assess the differences between LLM and textbook responses.

resultsA total of 90 surveys were completed. LLM responses were rated as more thorough (p < 0.001) but less concise (p < 0.001) than textbook responses. Textbooks were rated superior for answering questions on contraindications (p = 0.027) and complications (p = 0.014). ChatGPT was perceived as more accurate (p = 0.018), thorough (p = 0.002), and useful (p = 0.026) than Gemini. Gemini was rated lower in quality (p = 0.30) compared to ChatGPT along with being inferior to textbook answers for burn-related questions (p = 0.017) and anatomical questions (p = 0.013).

conclusionWhile LLMs show promise in generating thorough educational content, they require improvement in conciseness, accuracy, and utility for practical case preparation. ChatGPT generally outperforms Gemini, indicating variability in LLM capabilities. Further development should focus on enhancing accuracy and consistency to establish LLMs as reliable tools in medical education and practice.

Indexed as

LanguageSurgery, PlasticTextbooks as TopicClinical CompetenceFemaleHumansLarge Language ModelsMaleStudents, MedicalSurveys and QuestionnairesArtificial intelligenceChatgptEducationGeminiLarge language modelsMedical schoolPlastic and reconstructive surgeryResidencySurgery

Identifiers

PMID40597031
PMCPMC12220192

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.