Evidence map›Paper›PMID 42580688›Full record

ArticleJournal of medical Internet research2026

Critical Care-Specific vs General-Purpose Large Language Models in Emergency Intensive Care Unit Diagnosis: Single-Center Retrospective Paired Comparative Study.

Lihong Zheng, Zeyu Lin, Xiaolu Liu, Zhao Fan, Zhong He, Junjie Xu, Lu Yin

Abstract readComparative Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Lihong ZhengDepartment of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0009-0008-7593-6504
Zeyu LinDepartment of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0009-0002-9144-7846
Xiaolu LiuShenzhen University, Shenzhen, Guangdong, China.ORCID https://orcid.org/0009-0003-9954-3450
Zhao FanDepartment of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0009-0004-8495-8239
Zhong HeDepartment of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0009-0005-2239-2644
Junjie Xu *Clinical Research Institute, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0000-0003-4303-7295
Lu Yin *Department of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.ORCID https://orcid.org/0000-0001-9090-1299

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe emergency intensive care unit (EICU) manages the most critically ill patients, where rapid and accurate diagnosis is essential yet challenging. Diagnostic error rates in this setting are more than twice as high as in general wards, with serious consequences for patient outcomes. Large language models (LLMs) have attracted growing interest as decision-support tools; however, direct comparative evidence between critical care-specialized and general-purpose LLMs across the admission-to-discharge diagnostic workflow remains limited.

objectiveThis study aimed to compare the top-1 diagnostic accuracy of a critical care-specific LLM (Qiyuan 3.0.1) with 3 general-purpose models (GPT‑5.1, DeepSeek V3.1, and Qwen3‑32B) for EICU diseases, providing evidence for intelligent tool selection.

methodsThis single‑center retrospective paired study enrolled 184 consecutive EICU patients (April 2025-March 2026). Two standardized datasets were constructed: an initial dataset (first 24 hours of admission) and a final dataset (complete clinical course). All 4 models received identical zero‑shot prompts and generated diagnoses independently under masked conditions. The gold standard was the consensus diagnosis by 3 senior intensivists (>10 years' EICU experience; Fleiss κ=0.82). The primary end point was final‑stage top-1 accuracy; secondary end points were initial‑stage top-1 accuracy and the number of correctly matched diagnoses among the first 3 outputs at the final stage. Overall comparisons used Cochran Q test, followed by paired McNemar tests with Bonferroni correction; intergroup differences for top-3 counts were assessed by Friedman rank sum test.

resultsFinal-stage top-1 accuracy varied significantly across models (Cochran Q=20.32; P<.001): Qiyuan 3.0.1 reached 64.1% (118/184), followed by GPT-5.1 (109/184, 59.2%), DeepSeek V3.1 (105/184, 57.1%), and Qwen3-32B (95/184, 51.6%). Corrected pairwise comparisons (α=.0083) confirmed Qiyuan 3.0.1, GPT-5.1, and DeepSeek all outperformed Qwen3-32B significantly (all adjusted P<.008), while no statistical gaps were detected between Qiyuan 3.0.1 and the 2 top-performing general models. Though overall initial-stage accuracy differed significantly (Cochran Q=13.87; P<.001), no pairwise comparisons yielded significant results after correction. All models shared a median of 2 (IQR 1.0-2.0) correct top-3 diagnoses with no intergroup disparity (Friedman χ

conclusionsUnder the specific data conditions of this study, the critical care-specialized Qiyuan 3.0.1 performed comparably to leading general‑purpose LLMs (GPT‑5.1 and DeepSeek V3.1), supporting its potential for further specialty‑oriented exploration. Nevertheless, absolute accuracy below 70% precludes its direct deployment as an independent diagnostic standard. Bridging the gap from preliminary evaluation to clinical translation requires multicenter external validation, prospective human‑machine collaboration trials, and deeper model optimization-including sustained fine‑tuning on critical care corpora, transparent reasoning pathway design, and systematic safety boundary assessment.

Indexed as

Critical CareIntensive Care UnitsLarge Language ModelsAgedFemaleHumansMaleRetrospective Studiesartificial intelligence–assisted diagnosiscritical illnessdiagnostic accuracyEICUemergency intensive care unitlarge language modelsLLMs

Identifiers

PMID42580688
PMCPMC13507718

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.