Evidence map›Paper›PMID 42286747›Full record

ArticleSystematic reviews2026

Using full agreement across multiple large language models for title-and-abstract screening in systematic reviews: a proof-of-concept.

Frederic Hilkenmeier, Merle Stoltenberg, Christian Stierle

Abstract read
In one paragraph

Article in Systematic reviews, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Frederic HilkenmeierPsychology School, Hochschule Fresenius, Hamburg, Germany. Frederic.hilkenmeier@hs-fresenius.de.ORCID 0000-0002-5068-3108
Merle StoltenbergPsychology School, Hochschule Fresenius, Hamburg, Germany.
Christian StierlePsychology School, Hochschule Fresenius, Hamburg, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe exponential growth of scientific literature poses significant challenges for conducting systematic reviews, particularly in the labor-intensive title-and-abstract screening phase. This study examines the feasibility of using multiple large language models (LLMs) for title-and-abstract screening in systematic reviews.

methodsWe propose and evaluate a full-agreement approach using three commercially available LLMs from different model families (ChatGPT, Gemini, and Claude), in which automated classification decisions are only accepted when all three models assign the same label. This approach was examined across six datasets to assess its effectiveness. A structured workflow was developed to support implementation without requiring specialized technical expertise. Additionally, a stop criterion was introduced to ensure that LLM-based classification is only applied when predefined performance thresholds are met.

resultsAcross the six datasets, approximately four in five abstracts received full agreement across models and could therefore be classified automatically. For this subset of abstracts, classification performance was consistently higher than that of previous automated approaches using LLMs, with statistically significant improvements in all prespecified performance metrics.

conclusionsA full agreement approach across three LLMs may offer a promising and conservative strategy for title-and-abstract screening in systematic reviews by automating concordant decisions while reserving human review for discordant cases. As a proof-of-concept, the present findings support this approach as a possible workflow for reducing screening burden in the face of continued growth in the scientific literature, although its broader generalizability remains to be established.

Indexed as

Abstracting and IndexingLarge Language ModelsReview Literature as TopicSystematic Reviews as TopicHumansProof of Concept StudyAbstract classificationEvidence synthesisLarge language modelsMachine learningSystematic reviews

Identifiers

PMID42286747
PMCPMC13263951

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.