SynthesisNature medicine2026
LLM-assisted systematic review of large language models in clinical medicine.
Synthesis in Nature medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. It is linked to trial NCT07640906 (A Multicenter Comparative Study Evaluating the Impact of an AI-Assisted Chest CT Reporting System on Real-world Radiologist Performance), which is not on this map. Cited by 31 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
A Multicenter Comparative Study Evaluating the Impact of an AI-Assisted Chest CT Reporting System on Real-world Radiologist Performance: The DOUBLE-ACE Study
Who cites it
31 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review.Journal of medical Internet research · 2026Pooled it
- Application of Large Language Models in Chronic Disease Care: Mixed Methods Systematic Review and Thematic Synthesis.Journal of medical Internet research · 2026Pooled it
- Accountability for large language models in health care.Bulletin of the World Health Organization · 2026Article
- Article
- [Generative artificial intelligence and language models in trauma surgery : Applications in clinical care, research and teaching].Unfallchirurgie (Heidelberg, Germany) · 2026Review
- Prospective evidence for conversational medical AI is hard, but non-negotiable.Nature medicine · 2026Article
- Artificial Intelligence for Alzheimer's Disease Diagnosis: From Traditional Machine Learning to Large Language Models.Biosensors · 2026Review
- Exploring Bias in Medical Applications of Large Language Models: Protocol for a Systematic Review.JMIR research protocols · 2026Article
- Perspectives on the Limits and Clinical Alignment of Medical AI from Population Statistics to Individual Care.Bioengineering (Basel, Switzerland) · 2026Article
- Randomized Controlled Trial of Large Language Model-Assisted Diagnostic Accuracy in Nephrology.Kidney international reports · 2026Article
- Large language models in oncology: promise, pitfalls, and the path to real-world adoption.ESMO real world data and digital oncology · 2026Article
- Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis.JMIR AI · 2026Review
- The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.Journal of medical Internet research · 2026Article
- Evaluation Methods for Inference-Time Retrieval-Augmented and Graph Retrieval-Augmented Large Language Models in Health Care: Scoping Review.Journal of medical Internet research · 2026Article
- Development of an LLM pipeline exceeding physician-documented cardiovascular risk scores under routine clinical conditions.European heart journal. Digital health · 2026Article
- Applications of DeepSeek in Medicine: Bibliometric Analysis and Scoping Review.Journal of medical Internet research · 2026Article
- Large Language Models vs. Machine Learning on Structured Perioperative Data: Does Model Choice Matter?Journal of medical systems · 2026Article
- Mapping the AI life sciences landscape in Greece: a bibliometric comparison with global patterns.Scientific reports · 2026Article
- Automated full-text screening and accelerated reviews using large language models with context-aware agents: an exploratory analysis in biomarker research.European heart journal. Digital health · 2026Article
- Performance of ChatGPT-4o in Providing Information on Pediatric Inborn Errors of Immunity: A Cross-Sectional Evaluation.Journal of clinical medicine · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
12 authors.
Funding
Abstract
Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (n = 1,857) or exam-style tasks (n = 1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.