ArticleJAMA network open2026
Screening for Missed Opportunities for Diagnosis in the ED Using eTriggers and Large Language Models.
Article in JAMA network open, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
20 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Importance: Emergency department (ED) quality review often uses administrative electronic triggers (eTriggers), but yields on detecting missed opportunities for diagnosis (MODs) are low. A commercial large language model (LLM) may help screen for MODs, yet evaluation data in real-world cohorts remain limited. Objective: To evaluate LLMs for identifying MODs in ED eTrigger cohorts. Design, Setting, and Participants: This retrospective diagnostic study of 2 eTrigger cohorts, ED discharge with return hospital admission within 72 hours and ED admission to the floor with intensive care unit (ICU) escalation within 24 hours, was conducted from April 2015 through March 2025 across 9 EDs (2 academic and 7 community) in 1 US health system. Samples included 200 encounters from the 72-hour return cohort and 100 encounters from the floor-to-ICU cohort; each case was adjudicated by 2 emergency physicians using a review process based on the Safer Dx framework. Exposures: Cases were evaluated by Claude Sonnet 4, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Pro, GPT-5, and GPT-5 mini. Main Outcomes and Measures: Main outcomes were sensitivity, specificity, positive predictive value, negative predictive value, area under the receiver operating characteristic curve (AUC), and reviewer-reviewer and reviewer-model concordance. Results: Among 300 sampled encounters, 12 were excluded, leaving 288 analyzed encounters (median [IQR] age, 69 [54-79] years; 135 female [46.9%]) with 39 MODs (13.5%), including 21 of 191 (11.0%) in the 72-hour return cohort and 18 of 97 (18.6%) in the floor-to-ICU cohort. Interrater agreement was 81.9% (95% CI, 77.4%-86.1%), with Gwet AC1 of 0.77 (95% CI, 0.70-0.83). In the 72-hour return cohort, model sensitivity ranged from 42.9% (95% CI, 24.5%-63.5%) for GPT-5 mini to 85.7% (95% CI, 65.4%-95.0%) for Claude Sonnet 4, specificity from 55.9% (95% CI, 48.4%-63.1%) for Claude Sonnet 4 to 82.9% (95% CI, 76.6%-87.9%) for GPT-5 mini, and AUC from 0.65 (95% CI, 0.53-0.77) for GPT-5 mini to 0.73 (95% CI, 0.61-0.85) for Claude Sonnet 4. In the floor-to-ICU cohort, sensitivity ranged from 5.6% (95% CI, 1.0-25.8%) for GPT-5 mini to 55.6% (95% CI, 33.7%-75.4%) for Claude Sonnet 4, specificity from 64.6% (53.6%-74.2%) for Claude Sonnet 4 to 97.5% (95% CI, 91.2%-99.3%) for GPT-5 mini, and AUC from 0.57 (95% CI, 0.46-0.67) for GPT-5 mini to 0.82 (95% CI, 0.73-0.91) for GPT-5. Across cohorts, LLMs showed similar discrimination but different sensitivity-specificity tradeoffs; Claude Sonnet 4 generally favored higher sensitivity, whereas GPT-5 mini favored higher specificity. Conclusions and Relevance: In this diagnostic study of 2 ED eTrigger cohorts, model performance varied by cohort, with LLMs showing similar discrimination but different binary thresholds. These findings suggest that evaluation within the review workflow is needed before implementation and that reviewer-like concordance captures a distinct dimension of model behavior from discrimination.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.