Evidence map›Paper›PMID 42217753›Full record

Observational studyBrazilian journal of anesthesiology (Elsevier)

Evaluation of two large language models for intensive care unit discharge decisions: a prospective observational cohort study.

Engin İhsan Turan, Abdurrahman Engin Baydemir, Ebru Kaya, Zehra Polat Turan, Ayça Sultan Şahin

Registry-linked trialAbstract readObservational Study
In one paragraph

Observational study in Brazilian journal of anesthesiology (Elsevier). The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT06584890 (The Evaluation of the Effectiveness of General Artificial Intelligence Models in Extubation Decision-Making in the Intensive Care Unit), which is not on this map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT06584890 completednot on this map

The Evaluation of the Effectiveness of General Artificial Intelligence Models in Extubation Decision-Making in the Intensive Care Unit

TypeobservationalSponsorKanuni Sultan Suleyman Training and Research HospitalRan2024 to 2025Enrolled398ConditionsArtificial IntelegenceArmsdecision
3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Engin İhsan TuranIstanbul Health Science University Kanuni Sultan Süleyman Education and Training Hospital, Department of Anesthesiology, Istanbul, Turkey. Electronic address: enginihsan@hotmail.com.
Abdurrahman Engin BaydemirBasaksehir Cam ve Sakura City Hospital, Department of Anesthesiology, Istanbul, Turkey.
Ebru KayaIstanbul Health Science University Kanuni Sultan Süleyman Education and Training Hospital, Department of Anesthesiology, Istanbul, Turkey.
Zehra Polat TuranIstanbul Health Science University Sisli Hamidiye Etfal Education and Training Hospital, Department of Anesthesiology, Istanbul, Turkey.
Ayça Sultan ŞahinIstanbul Health Science University Kanuni Sultan Süleyman Education and Training Hospital, Department of Anesthesiology, Istanbul, Turkey.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe aim of this study was to evaluate the effectiveness of two general-purpose Large Language Models (LLMs), ChatGPT and Gemini, in predicting Intensive Care Unit (ICU) discharge decisions (discharge vs. non-discharge). By comparing their outputs with decisions made by ICU physicians, we sought to determine the alignment of AI-generated recommendations with expert clinical judgment and assess their potential as decision-support tools in critical care.

methodsThis prospective observational cohort study was conducted in a tertiary ICU between September 2024 and May 2025. Adult patients (≥ 18 years) requiring ICU discharge decisions were included. Standardized clinical prompts were generated from electronic health records and input into ChatGPT and Gemini. The models' binary discharge decisions were compared to those of ICU physicians. Model performance was assessed using accuracy, sensitivity, specificity, F1 score, Cohen's kappa, and McNemar's test. Discharge was defined as the positive class for all diagnostic performance analyses.

resultsA total of 398 patients were analyzed. ChatGPT demonstrated higher accuracy than Gemini (87.2% vs. 66.3%), with higher sensitivity (85.9% vs. 46.9%) and F1 score (0.890 vs. 0.628), whereas Gemini showed higher specificity (96.2% vs. 89.2%). Agreement with clinician decisions was substantial for ChatGPT (κ = 0.737, p = 0.024) and fair for Gemini (κ = 0.379, p < 0.001). Laboratory markers such as lactate, hemoglobin, and procalcitonin significantly differed between discharged and non-discharged patients.

conclusionLarge language models may support ICU discharge decisions when guided by structured, guideline-informed prompting. ChatGPT achieved higher overall accuracy, sensitivity, and F1 score, whereas Gemini demonstrated higher specificity.

trial registrationExternation (Discharge) of ICU, NCT06584890, registered 03 September 2024, prospectively registered, https://register. CLINICALTRIALS: gov/prs/beta/studies/S000EVXZ00000029/recordSummary.

Indexed as

Critical CareIntensive Care UnitsLarge Language ModelsPatient DischargeAdultAgedCohort StudiesFemaleGenerative Artificial IntelligenceHumansMaleMiddle AgedProspective StudiesSensitivity and SpecificityArtificial intelligenceDecision makingDecision support systems, ClinicalIntensive care unitsNatural language processingPatient discharge

Identifiers

PMID42217753
PMCPMC13312111

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.