SynthesisRegional anesthesia and pain medicine2026
Human versus artificial intelligence: evaluating ChatGPT's performance in conducting published systematic reviews with meta-analysis in chronic pain research.
Synthesis in Regional anesthesia and pain medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
8 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Agreement between ChatGPT and human-derived multilevel meta-analyses: a reproducibility study across clinical evidence syntheses.BMC medical research methodology · 2026Pooled it
- Moving towards acceleration with accountability: a conceptual framework for AI-assisted systematic reviews.Global epidemiology · 2026Article
- Artificial Intelligence Resources for the Screening of Titles and Abstracts in Systematic Reviews: A Scoping Review.Cochrane evidence synthesis and methods · 2026Review
- Large language models are comparable with commonly used statistical software: A validation of GPT 5.1 for frequentist meta-analysis in orthopaedics.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Article
- The Influencing Factors of Medical Postgraduates' Usage Intention Toward Artificial Intelligence-Generated Content Tools in Academic Research: Qualitative Analysis.Journal of medical Internet research · 2026Article
- Efficacy of Large Language Models for Screening of Systematic Reviews on Periprosthetic Joint Infection.Journal of clinical medicine · 2026Article
- The future of reviews : Will LLMs render them obsolete?EMBO reports · 2025Article
- Multiple Confabulations Found in Bioinformatics Tasks Carried Out by Several Free Large Language Models.Current genomics · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introductionArtificial intelligence (AI), particularly large-language models like Chat Generative Pre-Trained Transformer (ChatGPT), has demonstrated potential in streamlining research methodologies. Systematic reviews and meta-analyses, often considered the pinnacle of evidence-based medicine, are inherently time-intensive and demand meticulous planning, rigorous data extraction, thorough analysis, and careful synthesis. Despite promising applications of AI, its utility in conducting systematic reviews with meta-analysis remains unclear. This study evaluated ChatGPT's accuracy in conducting key tasks of a systematic review with meta-analysis.
methodsThis validation study used data from a published meta-analysis on emotional functioning after spinal cord stimulation. ChatGPT-4o performed title/abstract screening, full-text study selection, and data pooling for this systematic review with meta-analysis. Comparisons were made against human-executed steps, which were considered the gold standard. Outcomes of interest included accuracy, sensitivity, specificity, positive predictive value, and negative predictive value for screening and full-text review tasks. We also assessed for discrepancies in pooled effect estimates and forest plot generation.
resultsFor title and abstract screening, ChatGPT achieved an accuracy of 70.4%, sensitivity of 54.9%, and specificity of 80.1%. In the full-text screening phase, accuracy was 68.4%, sensitivity 75.6%, and specificity 66.8%. ChatGPT successfully pooled data for five forest plots, achieving 100% accuracy in calculating pooled mean differences, 95% CIs, and heterogeneity estimates (
conclusionChatGPT demonstrates modest to moderate accuracy in screening and study selection tasks, but performs well in data pooling and meta-analytic calculations. These findings underscore the potential of AI to augment systematic review methodologies, while also emphasizing the need for human oversight to ensure accuracy and integrity in research workflows.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.