Evidence map›Paper›PMID 42531250›Full record

ArticleJournal of medical Internet research2026

A Multiagent Large Language Model Framework for Emergency Treatment Recommendation in Acute Ischemic Stroke: Development and Validation Study.

Bicong Yan, Ruipeng Zhang, Li Chen, Xinyu Song, Zhongzheng Cao, Yuehua Li

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Bicong Yan *Department of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0000-0002-2003-3120
Ruipeng Zhang *Department of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0000-0002-4372-4987
Li ChenDepartment of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0000-0001-9971-9653
Xinyu SongDepartment of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0000-0002-3539-2492
Zhongzheng CaoDepartment of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0009-0006-7885-2467
Yuehua LiDepartment of Diagnostic and Interventional Radiology, Shanghai Sixth People's Hospital, No.600 Yishan Road, Xuhui District, Shanghai, 200233, China, 86 18918727305.ORCID http://orcid.org/0000-0001-8028-248X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Acute ischemic stroke (AIS) treatment selection requires rapid, guideline-concordant integration of clinical, imaging, and laboratory data, including therapeutic windows, contraindications, stroke severity, and imaging eligibility. This process is complex, expertise-dependent, and vulnerable to safety-critical errors. Objective: This study aimed to develop and validate a structured multiagent large language model (LLM) framework for AIS decision support using real-world cases and to assess its accuracy, safety, auditability, and impact on physician decision-making, particularly among junior physicians and nonspecialists. Methods: We developed a multiagent LLM framework that used structured outputs and guideline-based reasoning to generate treatment recommendations (intravenous thrombolysis, endovascular thrombectomy, standard medical therapy, or non-AIS, or nonstroke) and Trial of ORG 10172 in Acute Stroke Treatment (TOAST) classification. The framework was evaluated using multicenter retrospective real-world cases from 2 hospitals collected between January 2018 and March 2025, prospective cases from February to May 2025, and literature-derived challenging cases from PubMed between January 2024 and January 2025. Performance was assessed against clinical reference standards. Safety was assessed using omission and hallucination event rates, instruction adherence, and 5-point clinical safety ratings. In a prospective physician study, physicians with different seniority and specialty backgrounds made AIS treatment and TOAST classification decisions with and without LLM support. Physician-case decision-level outcomes were analyzed using a binomial generalized linear mixed-effects model accounting for physician and case effects. Results: The final analysis included 1055 group A cases, 721 group B cases, 144 literature-derived group C cases, and 161 prospectively collected group D cases. Across representative Baichuan, Qwen, DeepSeek, and GPT models, the multiagent framework consistently improved treatment recommendation accuracy. Model-level accuracy ranges increased from 0.546-0.737 to 0.687-0.851 in group A, from 0.587-0.698 to 0.671-0.813 in group B, and from 0.507-0.646 to 0.667-0.750 in group C. TOAST classification improved overall, with cohort-level variation. Across evaluated models, the multiagent framework increased the mean clinical safety score from 3.70 to 4.01 and reduced mean hallucination and omission rates from 33.6% to 20.6% and from 38.5% to 24.5%, respectively. In the prospective physician study, LLM support increased treatment decision accuracy from 73.1% to 88.6% (odds ratio 2.86, 95% CI, 2.27-3.60; P<.001). Accuracy gains were largest among junior and nonspecialist physicians, including junior specialists (0.667 to 0.833), junior nonspecialists (0.600 to 0.846), and senior nonspecialists (0.667 to 0.850). TOAST classification performance also improved (odds ratio 3.63, 95% CI 2.85-4.64; P<.001). Conclusions: A structured multiagent framework improved LLM performance with average improvements of 18.9% in AIS treatment recommendation and TOAST classification, while producing more structured, auditable outputs with higher safety ratings. It was associated with higher physician decision accuracy, with larger gains among less-experienced physicians, suggesting the potential to narrow expertise-related decision accuracy gaps. Prospective multicenter studies are needed to assess effects on workflow and clinical outcomes.

Indexed as

Decision Support Systems, ClinicalIschemic StrokeStrokeHumansLarge Language ModelsRetrospective Studiesacute ischemic strokeclinical safetydecision supportlarge language modelmultiagent system

Identifiers

PMID42531250
PMCPMC13423760

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.