ArticleBMC medical research methodology2020
An evaluation of DistillerSR's machine learning-based prioritization tool for title/abstract screening - impact on reviewer-relevant outcomes.
Article in BMC medical research methodology, 2020. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 65 papers, 5 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
65 citing papers in PubMed, 5 syntheses or guidelines pooled it.
- Systematic evidence map on the association between exposure to personal care products and fetal growth.Environment international · 2026Pooled it
- A systematic review and meta-analysis of observational studies and uncontrolled trials reporting on the use of checkpoint blockers in patients with cancer and pre-existing autoimmune disease.European journal of cancer (Oxford, England : 1990) · 2024Pooled it
- Estimated global and regional causes of deaths from diarrhoea in children younger than 5 years during 2000-21: a systematic review and Bayesian multinomial analysis.The Lancet. Global health · 2024Pooled it
- Patient preferences for breast cancer screening: a systematic review update to inform recommendations by the Canadian Task Force on Preventive Health Care.Systematic reviews · 2024Pooled it
- Identifying a list of healthcare 'never events' to effect system change: a systematic review and narrative synthesis.BMJ open quality · 2023Pooled it
- Developing and Evaluating the Use of ChatGPT as a Screening Tool for Nurses Conducting Structured Literature Reviews: Proof of Concept Study Results.Journal of clinical nursing · 2026Article
- Techniques, Performance, and Feasibility of Natural Language Processing for Abstract Screening in Evidence Synthesis: A Systematic Review.Campbell systematic reviews · 2026Review
- Artificial Intelligence Resources for the Screening of Titles and Abstracts in Systematic Reviews: A Scoping Review.Cochrane evidence synthesis and methods · 2026Review
- Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study.JMIR formative research · 2026Article
- Performance of Two AI Approaches in ASReview Compared With Manual Screening for Dementia Care Literature Screening: Comparative Analysis.JMIR formative research · 2026Article
- Connecting Patients with Clinical Trials Using Patient Navigation: A Scoping Review.Current oncology (Toronto, Ont.) · 2026Article
- Article
- Comparing the performance of narrow vs. broad search strategies when using machine learning-based software for title/abstract screening.Journal of the Medical Library Association : JMLA · 2026Article
- Screenathon 2.0: human-AI collaborative screening applied to patient-generated health data.Scientific reports · 2026Article
- The landscape of artificial intelligence tools and platforms for evidence synthesis: a scoping review.Systematic reviews · 2026Article
- Detecting false exclusions in single-reviewer literature screening by using AI tools as secondary reviewers: a study protocol for an evaluation study.Systematic reviews · 2026Article
- Health effects associated with alcohol consumption: a Burden of Proof study.Nature health · 2026Article
- Review
- The hunt for the last relevant paper: blending the best of humans and AI.European journal of psychotraumatology · 2025Article
- Protocol: Understanding the Content, Context, and Impact of Far-Right Extremist Propaganda Disseminated Online: A Systematic Review.Campbell systematic reviews · 2025Article
5 more citing papers are in PubMed but not listed here.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
Abstract
backgroundSystematic reviews often require substantial resources, partially due to the large number of records identified during searching. Although artificial intelligence may not be ready to fully replace human reviewers, it may accelerate and reduce the screening burden. Using DistillerSR (May 2020 release), we evaluated the performance of the prioritization simulation tool to determine the reduction in screening burden and time savings.
methodsUsing a true recall @ 95%, response sets from 10 completed systematic reviews were used to evaluate: (i) the reduction of screening burden; (ii) the accuracy of the prioritization algorithm; and (iii) the hours saved when a modified screening approach was implemented. To account for variation in the simulations, and to introduce randomness (through shuffling the references), 10 simulations were run for each review. Means, standard deviations, medians and interquartile ranges (IQR) are presented.
resultsAmong the 10 systematic reviews, using true recall @ 95% there was a median reduction in screening burden of 47.1% (IQR: 37.5 to 58.0%). A median of 41.2% (IQR: 33.4 to 46.9%) of the excluded records needed to be screened to achieve true recall @ 95%. The median title/abstract screening hours saved using a modified screening approach at a true recall @ 95% was 29.8 h (IQR: 28.1 to 74.7 h). This was increased to a median of 36 h (IQR: 32.2 to 79.7 h) when considering the time saved not retrieving and screening full texts of the remaining 5% of records not yet identified as included at title/abstract. Among the 100 simulations (10 simulations per review), none of these 5% of records were a final included study in the systematic review. The reduction in screening burden to achieve true recall @ 95% compared to @ 100% resulted in a reduced screening burden median of 40.6% (IQR: 38.3 to 54.2%).
conclusionsThe prioritization tool in DistillerSR can reduce screening burden. A modified or stop screening approach once a true recall @ 95% is achieved appears to be a valid method for rapid reviews, and perhaps systematic reviews. This needs to be further evaluated in prospective reviews using the estimated recall.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.