Evidence map›Paper›PMID 41645231›Full record

ArticleScandinavian journal of trauma, resuscitation and emergency medicine2026

Multimodal large language model versus emergency physicians for burn assessment: a prospective non-inferiority study.

Ahmet Aykut, Ali Rıza Karayıl, Cem Yıldırım, Ertuğ Günsoy, Mehmet Tatlı, Murat Avcı

Abstract readEquivalence Trial
In one paragraph

Article in Scandinavian journal of trauma, resuscitation and emergency medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Ahmet AykutDepartment of Emergency Medicine, SBU Van Education and Research Hospital, Van, Türkiye. ahmet.aykut@gmail.com.ORCID http://orcid.org/0009-0001-3173-8994
Ali Rıza KarayılDepartment of General Surgery, SBU Van Education and Research Hospital, Van, Türkiye.ORCID http://orcid.org/0000-0002-4080-4455
Cem YıldırımDepartment of Emergency Medicine, SBU Van Education and Research Hospital, Van, Türkiye.ORCID http://orcid.org/0009-0004-5359-6335
Ertuğ GünsoyDepartment of Emergency Medicine, SBU Van Education and Research Hospital, Van, Türkiye.ORCID http://orcid.org/0000-0002-9653-3236
Mehmet TatlıDepartment of Emergency Medicine, SBU Van Education and Research Hospital, Van, Türkiye.ORCID http://orcid.org/0000-0001-5907-9161
Murat AvcıDepartment of Emergency Medicine, SBU Van Education and Research Hospital, Van, Türkiye.ORCID http://orcid.org/0009-0007-0025-1108

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAccurate burn size and depth assessment at first contact guides fluid resuscitation, referral, and operative planning, yet both tasks show meaningful inter-clinician variability. General-purpose multimodal large language models may offer scalable, image-based decision support in emergency care, but prospective benchmarking against clinicians and a robust reference standard remains limited.

methodsWe conducted a prospective, single-centre diagnostic accuracy and agreement study in a tertiary emergency department (22 July-8 September 2025). Consecutive acute burn presentations (< 24 h) were screened; protocol-conformant cases contributed standardized three-view photographs per anatomically distinct burn region. A multimodal large language model generated region-level estimates of total body surface area (TBSA) contribution and burn depth class. Eighteen emergency physicians independently rated the same images and minimal metadata, blinded to model and reference outputs. A three-member expert panel served as the reference standard by consensus. The primary endpoint was non-inferiority of the model versus the physician median for region-level absolute TBSA error relative to the panel, with a pre-specified margin of 3 percentage points, using patient-level cluster bootstrap for inference. Secondary endpoints included TBSA agreement and depth agreement (quadratic-weighted kappa).

resultsOf 413 screened presentations, 52 patients were enrolled, yielding 64 analyzable burn region-cases (35 pediatric, 29 adult). The model's mean absolute TBSA error versus the panel was 1.40 percentage points (median 1.00); 87.5% of cases were within ± 3 percentage points and 98.4% within ± 5. The physician median had a mean absolute error of 0.89 percentage points (median 0.75). The paired non-inferiority analysis met the pre-specified criterion (Hodges-Lehmann median Δ = 0.25; one-sided 95% upper bound = 0.50), indicating the model was non-inferior to physicians for TBSA estimation. In contrast, depth agreement versus the panel was slight for the model (quadratic-weighted kappa 0.14), with systematic underestimation of deeper burns, while physician consensus showed substantially higher agreement (quadratic-weighted kappa 0.65).

conclusionsIn this prospective emergency department evaluation, a general-purpose multimodal model achieved non-inferior performance to emergency physicians for region-level TBSA estimation but performed substantially worse for burn depth classification. These findings support a narrowly defined adjunct role for TBSA estimation, while depth-dependent decisions should remain clinician-led and require further method development and external validation.

Indexed as

BurnsPhysiciansAdultBody Surface AreaChildEmergency Service, HospitalFemaleHumansLarge Language ModelsMaleMiddle AgedProspective StudiesBurnsDiagnostic accuracyEmergency medicineLarge language modelTotal body surface area

Identifiers

PMID41645231
PMCPMC12969848

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.