Evidence map›Paper›PMID 39956557›Full record

SynthesisRegional anesthesia and pain medicine2026

Human versus artificial intelligence: evaluating ChatGPT's performance in conducting published systematic reviews with meta-analysis in chronic pain research.

Anam Purewal, Kalli Fautsch, Johana Klasova, Nasir Hussain, Ryan S D'Souza

Abstract readComparative StudyMeta-AnalysisValidation Study
In one paragraph

Synthesis in Regional anesthesia and pain medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Review
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Anam PurewalDepartment of Orthopedic Surgery and Rehabilitation Medicine, SUNY Downstate Health Sciences University, Brooklyn, New York, USA.ORCID 0000-0002-3674-9019
Kalli FautschDepartment of Anesthesiology and Perioperative Medicine, Mayo Clinic, Rochester, Minnesota, USA.
Johana KlasovaDepartment of Anesthesiology and Perioperative Medicine, Mayo Clinic, Rochester, Minnesota, USA.
Nasir HussainThe Ohio State University, Columbus, Ohio, USA.ORCID 0000-0003-0353-1002
Ryan S D'SouzaDepartment of Anesthesiology and Perioperative Medicine, Mayo Clinic, Rochester, Minnesota, USA DSouza.Ryan@mayo.edu.ORCID 0000-0002-4601-9837

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionArtificial intelligence (AI), particularly large-language models like Chat Generative Pre-Trained Transformer (ChatGPT), has demonstrated potential in streamlining research methodologies. Systematic reviews and meta-analyses, often considered the pinnacle of evidence-based medicine, are inherently time-intensive and demand meticulous planning, rigorous data extraction, thorough analysis, and careful synthesis. Despite promising applications of AI, its utility in conducting systematic reviews with meta-analysis remains unclear. This study evaluated ChatGPT's accuracy in conducting key tasks of a systematic review with meta-analysis.

methodsThis validation study used data from a published meta-analysis on emotional functioning after spinal cord stimulation. ChatGPT-4o performed title/abstract screening, full-text study selection, and data pooling for this systematic review with meta-analysis. Comparisons were made against human-executed steps, which were considered the gold standard. Outcomes of interest included accuracy, sensitivity, specificity, positive predictive value, and negative predictive value for screening and full-text review tasks. We also assessed for discrepancies in pooled effect estimates and forest plot generation.

resultsFor title and abstract screening, ChatGPT achieved an accuracy of 70.4%, sensitivity of 54.9%, and specificity of 80.1%. In the full-text screening phase, accuracy was 68.4%, sensitivity 75.6%, and specificity 66.8%. ChatGPT successfully pooled data for five forest plots, achieving 100% accuracy in calculating pooled mean differences, 95% CIs, and heterogeneity estimates (

conclusionChatGPT demonstrates modest to moderate accuracy in screening and study selection tasks, but performs well in data pooling and meta-analytic calculations. These findings underscore the potential of AI to augment systematic review methodologies, while also emphasizing the need for human oversight to ensure accuracy and integrity in research workflows.

Indexed as

Artificial IntelligenceChronic PainMeta-Analysis as TopicSystematic Reviews as TopicGenerative Artificial IntelligenceHumansCHRONIC PAINMeta-AnalysisMethodsSpinal Cord Stimulation

Identifiers

PMID39956557
PMCPMC13151432

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.