Evidence map›Paper›PMID 42106619›Full record

ArticleBMC medical research methodology2026

When to stop reviewing: validation of stop criteria in ASReview.

C Kempny, K Annac, D Wahidie, Y Yilmaz-Aslan, P Brzoska

Abstract read
In one paragraph

Article in BMC medical research methodology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

C KempnyHealth Services Research, Witten/Herdecke University, Faculty of Health, School of Medicine, Witten, Germany. christian.kempny@uni-wh.de.
K AnnacHealth Services Research, Witten/Herdecke University, Faculty of Health, School of Medicine, Witten, Germany.
D WahidieHealth Services Research, Witten/Herdecke University, Faculty of Health, School of Medicine, Witten, Germany.
Y Yilmaz-AslanHealth Services Research, Witten/Herdecke University, Faculty of Health, School of Medicine, Witten, Germany.
P BrzoskaHealth Services Research, Witten/Herdecke University, Faculty of Health, School of Medicine, Witten, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundSystematic reviews are essential for evidence-based research, but they often require a great deal of time and effort. Although title and abstract (T&A) screening is just one part of the review process, it can be very time-consuming when search strategies retrieve large numbers of records. Given the exponential growth of scientific publications in recent decades, tools such as ASReview, which use machine learning (ML) for active learning-based screening, aim to reduce the workload. However, since ASReview helps users prioritise potentially relevant studies rather than supporting them in screening all records, a key question arises: at what point can the screening process safely stop without risking the omission of relevant studies when the entire dataset is not being reviewed?

methodsThis simulation study tested three proposed stop criteria for terminating screening in ASReview without loss of relevant data: (1) stopping after a calculated number of relevant studies based on an initial sample; (2) stopping after a fixed number of consecutively studies deemed irrelevant; (3) stopping after a predefined percentage of the dataset has been screened. A total of 35,000 automated title and abstract screenings were conducted using five datasets from the SYNERGY repository. Key outcomes included the percentage of studies screened until the last relevant study was found and the number of relevant studies missed under each stop criterion.

resultsThe proportion of the dataset that needed to be screened to identify all relevant studies (as pre-classified in the SYNERGY dataset) varied greatly across datasets, ranging from 2.9% to 76.9% on average. None of the tested stop criteria could consistently identify all relevant studies across all datasets. Stop criterion 1 was reliable in only 2% of simulations. Stop criterion 2 showed high variability, with thresholds ranging from 2% to 61%, depending on the dataset. Stop criterion 3 failed to define a universal percentage applicable across datasets.

conclusionsASReview can reduce screening workload by prioritizing potentially relevant studies through ML-based ranking, thereby allowing researchers to identify relevant studies earlier in the screening process. However, no stop criterion reliably ensures that all relevant studies are identified. Early stopping may result in missed studies, depending on dataset characteristics. Current stop criteria should be applied cautiously and potentially combined with quality assurance measures. Further research is needed to develop more robust and generalizable stopping rules.

trial registrationNot applicable - this is a simulation study, not a registered systematic review.

Indexed as

Evidence-Based MedicineMachine LearningReview Literature as TopicComputer SimulationHumansASReviewMachine learningScreening efficiencyStop CriteriaSystematic reviewsTitle and abstract screening

Identifiers

PMID42106619
PMCPMC13156873

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.