ArticleCochrane evidence synthesis and methods2025
Using a Large Language Model (ChatGPT-4o) to Assess the Risk of Bias in Randomized Controlled Trials of Medical Interventions: Interrater Agreement With Human Reviewers.
Article in Cochrane evidence synthesis and methods, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
4 citing papers in PubMed.
- Article
- Evidence-Based Medicine Meets Artificial Intelligence: Reframing Critical Appraisal in Emergency Medicine.Academic emergency medicine : official journal of the Society for Academic Emergency Medicine · 2026Article
- Evidence-to-Decision Frameworks: Enhancing the Quality and Rigour of Guidelines and Recommendations.Clinical and public health guidelines · 2026Article
- Comparison of three large language models' ability to assess the risk of bias using ROBINS-I tool.BMJ digital health & AI · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Risk of bias (RoB) assessment is a highly skilled task that is time-consuming and subject to human error. RoB automation tools have previously used machine learning models built using relatively small task-specific training sets. Large language models (LLMs; e.g., ChatGPT) are complex models built using non-task-specific Internet-scale training sets. They demonstrate human-like abilities and might be able to support tasks like RoB assessment. Methods: Following a published peer-reviewed protocol, we randomly sampled 100 Cochrane reviews. New or updated reviews that evaluated medical interventions, included ≥ 1 eligible trial, and presented human consensus assessments using Cochrane RoB1 or RoB2 were eligible. We excluded reviews performed under emergency conditions (e.g., COVID-19), and those on public health or welfare. We randomly sampled one trial from each review. Trials using individual- or cluster-randomized designs were eligible. We extracted human consensus RoB assessments of the trials from the reviews, and methods texts from the trials. We used 25 review-trial pairs to develop a ChatGPT prompt to assess RoB using trial methods text. We used the prompt and the remaining 75 review-trial pairs to estimate human-ChatGPT agreement for "Overall RoB" (primary outcome) and "RoB due to the randomization process", and ChatGPT-ChatGPT (intrarater) agreement for "Overall RoB". We used ChatGPT-4o (February 2025) throughout. Results: The 75 reviews were sampled from 35 Cochrane review groups, and all used RoB1. The 75 trials spanned five decades, and all but one were published in English. Human-ChatGPT agreement for "Overall RoB" assessment was 50.7% (95% CI 39.3%-62.0%), substantially higher than expected by chance ( Conclusions: ChatGPT appears to have some ability to assess RoB and is unlikely to be guessing or "hallucinating". The estimated agreement for "Overall RoB" is well above estimates of agreement reported for some human reviewers, but below the highest estimates. LLM-based systems for assessing RoB may be able to help streamline and improve evidence synthesis production.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.