Evidence map›Paper›PMID 42423173›Full record

ArticleWorld journal of surgery2026

Bedside Triage by Large Language Models in Acute Pancreatitis: A Scenario-Based Comparative Evaluation of GPT-4, GPT-5, and Gemini.

Yahya Kemal Çalışkan, Fatih Başak, Olgun Erdem

Abstract readComparative Study
In one paragraph

Article in World journal of surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Yahya Kemal ÇalışkanDepartment of General Surgery, University of Health Sciences, Kanuni Training and Research Hospital, Istanbul, Turkey.ORCID https://orcid.org/0000-0003-1999-1601
Fatih BaşakDepartment of General Surgery, University of Health Sciences, Umraniye Training and Research Hospital, Istanbul, Turkey.ORCID https://orcid.org/0000-0003-1854-7437
Olgun ErdemDepartment of General Surgery, University of Health Sciences, Umraniye Training and Research Hospital, Istanbul, Turkey.ORCID https://orcid.org/0000-0003-1433-3431

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundEarly decision-making in acute pancreatitis (AP) involves diagnostic confirmation, early severity triage, escalation thresholds, and initiation of guideline-concordant management under time pressure and incomplete information. Large language models (LLMs) may support structured bedside reasoning, but their clinical usefulness cannot be inferred from guideline knowledge alone.

methodsA cross-sectional, scenario-based comparative evaluation was conducted in January 2026 using 20 AP scenarios: 15 refined hypothetical vignettes and 5 de-identified, privacy-modified real-life case patterns. GPT-4, GPT-5, and Gemini received identical single-turn prompts. Model access was through OpenAI API gpt-4-0613, OpenAI API gpt-5, and Google Vertex AI Gemini 1.0 Pro; temperature was set to 0.0, and each prompt was repeated three times per model. Outputs were scored by two independent clinician-raters using a prespecified 1-5 ordinal rubric across guideline concordance, safety, actionability, and data-synthesis quality. Two senior board-certified surgeons independently generated expert reference pathways for comparison.

resultsGPT-5 achieved the highest guideline concordance (4.28 ± 0.38) and safety (4.20 ± 0.45) profiles. GPT-4 provided the clearest stepwise actionability (4.15 ± 0.48), whereas Gemini showed the strongest data-synthesis quality (4.22 ± 0.52). With deterministic settings, internal consistency across three repeated runs was 100%. All models demonstrated clinically relevant failure modes, particularly unwarranted certainty under missing data; this occurred in 12/20 GPT-4, 7/20 GPT-5, and 15/20 Gemini outputs.

conclusionNo model should be used as a stand-alone bedside decision-maker for AP. In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.

Indexed as

Clinical Decision-MakingLarge Language ModelsPancreatitisTriageAcute DiseaseCross-Sectional StudiesGenerative Artificial IntelligenceHumansacute pancreatitisclinical decision supportguideline concordancelarge language modelssafetyscenario‐based evaluationtriage

Identifiers

PMID42423173
PMCPMC13460903

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.