Evidence map›Paper›PMID 40969781›Full record

ArticleCochrane evidence synthesis and methods2025

Using a Large Language Model (ChatGPT-4o) to Assess the Risk of Bias in Randomized Controlled Trials of Medical Interventions: Interrater Agreement With Human Reviewers.

Christopher James Rose, Julia Bidonde, Martin Ringsten, Julie Glanville, Thomas Potrebny, Chris Cooper, Ashley Elizabeth Muller, Hans Bugge Bergsund, Jose F Meneses-Echavez, Rigmor C Berg

Abstract read
In one paragraph

Article in Cochrane evidence synthesis and methods, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Evidence-Based Medicine Meets Artificial Intelligence: Reframing Critical Appraisal in Emergency Medicine.Academic emergency medicine : official journal of the Society for Academic Emergency Medicine · 2026
    Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Christopher James RoseCenter for Epidemic Interventions Research Norwegian Institute of Public Health Oslo Norway.ORCID https://orcid.org/0000-0001-6457-8168
Julia BidondeDivision of Health Services Norwegian Institute of Public Health Oslo Norway.
Martin RingstenCochrane Sweden Skåne University Hospital, Lund University Lund Sweden.ORCID https://orcid.org/0000-0003-4456-9760
Julie GlanvilleGlanville.info York UK.
Thomas PotrebnySection Evidence-Based Practice Western Norway University of Applied Sciences Bergen Norway.
Chris CooperBristol Medical School University of Bristol Bristol UK.ORCID https://orcid.org/0000-0003-0864-5607
Ashley Elizabeth MullerDivision of Health Services Norwegian Institute of Public Health Oslo Norway.
Hans Bugge BergsundDivision of Health Services Norwegian Institute of Public Health Oslo Norway.
Jose F Meneses-EchavezDivision of Health Services Norwegian Institute of Public Health Oslo Norway.
Rigmor C BergDivision of Health Services Norwegian Institute of Public Health Oslo Norway.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Risk of bias (RoB) assessment is a highly skilled task that is time-consuming and subject to human error. RoB automation tools have previously used machine learning models built using relatively small task-specific training sets. Large language models (LLMs; e.g., ChatGPT) are complex models built using non-task-specific Internet-scale training sets. They demonstrate human-like abilities and might be able to support tasks like RoB assessment. Methods: Following a published peer-reviewed protocol, we randomly sampled 100 Cochrane reviews. New or updated reviews that evaluated medical interventions, included ≥ 1 eligible trial, and presented human consensus assessments using Cochrane RoB1 or RoB2 were eligible. We excluded reviews performed under emergency conditions (e.g., COVID-19), and those on public health or welfare. We randomly sampled one trial from each review. Trials using individual- or cluster-randomized designs were eligible. We extracted human consensus RoB assessments of the trials from the reviews, and methods texts from the trials. We used 25 review-trial pairs to develop a ChatGPT prompt to assess RoB using trial methods text. We used the prompt and the remaining 75 review-trial pairs to estimate human-ChatGPT agreement for "Overall RoB" (primary outcome) and "RoB due to the randomization process", and ChatGPT-ChatGPT (intrarater) agreement for "Overall RoB". We used ChatGPT-4o (February 2025) throughout. Results: The 75 reviews were sampled from 35 Cochrane review groups, and all used RoB1. The 75 trials spanned five decades, and all but one were published in English. Human-ChatGPT agreement for "Overall RoB" assessment was 50.7% (95% CI 39.3%-62.0%), substantially higher than expected by chance ( Conclusions: ChatGPT appears to have some ability to assess RoB and is unlikely to be guessing or "hallucinating". The estimated agreement for "Overall RoB" is well above estimates of agreement reported for some human reviewers, but below the highest estimates. LLM-based systems for assessing RoB may be able to help streamline and improve evidence synthesis production.

Indexed as

artificial intelligenceChatGPTevidence synthesislarge language modelLLMrisk of biasRoB

Identifiers

PMID40969781
PMCPMC12442625

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.