Evidence map›Paper›PMID 41776077›Full record

SynthesisNature medicine2026

LLM-assisted systematic review of large language models in clinical medicine.

Sully F Chen, Anton Alyakin, Andreas Seas, Eunice Yang, Joanne J Choi, Jin Vivian Lee, Amelia L Chen, Pranav I Warman, Rochelle T Bitolas, Robert J Steele and 2 more

Registry-linked trialAbstract readSystematic Review
In one paragraph

Synthesis in Nature medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07640906 (A Multicenter Comparative Study Evaluating the Impact of an AI-Assisted Chest CT Reporting System on Real-world Radiologist Performance), which is not on this map. Cited by 31 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
31citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

NCT07640906 not yet recruitingnot on this map

A Multicenter Comparative Study Evaluating the Impact of an AI-Assisted Chest CT Reporting System on Real-world Radiologist Performance: The DOUBLE-ACE Study

TypeobservationalSponsorShanghai Zhongshan HospitalRan2026 to 2026Enrolled75ConditionsThoracic Diseases, Chest CT Scan, Artificial Intelligence (AI) in DiagnosisArmsAn AI-assisted reporting system integrated into the clinical workflow, providing automated draft generation to assist with chest CT interpretation
3 · Its place in the literature

Who cites it

31 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Accountability for large language models in health care.Bulletin of the World Health Organization · 2026
    Article
  4. Article
  5. Review
  6. Article
  7. Review
  8. Article
  9. Article
  10. Article
  11. Article
  12. Review
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Sully F ChenDuke University School of Medicine, Durham, NC, USA. sully.chen@duke.edu.ORCID http://orcid.org/0000-0001-7719-469X
Anton AlyakinWashington University School of Medicine, Saint Louis, MO, USA.
Andreas SeasDuke University School of Medicine, Durham, NC, USA.ORCID http://orcid.org/0000-0003-0624-1254
Eunice YangDepartment of Neurosurgery, NYU Langone Health, New York, NY, USA.
Joanne J ChoiDepartment of Neurosurgery, NYU Langone Health, New York, NY, USA.
Jin Vivian LeeWashington University School of Medicine, Saint Louis, MO, USA.ORCID http://orcid.org/0000-0002-8961-0125
Amelia L ChenUQ Ochsner, New Orleans, LA, USA.
Pranav I WarmanDepartment of Neurosurgery, Johns Hopkins University, Baltimore, MD, USA.
Rochelle T BitolasDuke University School of Medicine, Durham, NC, USA.ORCID http://orcid.org/0009-0002-3662-6745
Robert J SteeleDepartment of Neurosurgery, NYU Langone Health, New York, NY, USA.
Daniel A AlberNew York University Grossman School of Medicine, New York, NY, USA.ORCID http://orcid.org/0000-0001-7957-5170
Eric K OermannDepartment of Neurosurgery, NYU Langone Health, New York, NY, USA. eric.oermann@nyulangone.org.ORCID http://orcid.org/0000-0002-1876-5963

Funding

Medical Scientist Training Program Training GrantT32GM145449 · NIGMS · DUKE UNIVERSITY · PI Christopher D Kontos · 2022 to 2026
$6.6M
NIGMS NIH HHS T32 GM145449
6 · The paper itself

Abstract

Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (n = 1,857) or exam-style tasks (n = 1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption.

Indexed as

Large Language ModelsGenerative Artificial IntelligenceHumans

Identifiers

PMID41776077
PMCPMC13004689

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.