Evidence map›Paper›PMID 42059529›Full record

ArticleCanadian journal of psychiatry. Revue canadienne de psychiatrie2026

Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in Psychiatry: Nouvel algorithme de sélection des résumés utilisant un consensus issu d'un grand modèle de langage inspiré de la méthode Delphi pour les revues systématiques en psychiatrie.

Mirkamal Tolend, Ramzi Halabi, Kousai Ghaouari, Yvonne C Y Lau, Martin Alda, Arend Hintze, Benoit H Mulsant, Abigail Ortiz

Abstract read
In one paragraph

Article in Canadian journal of psychiatry. Revue canadienne de psychiatrie, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Mirkamal TolendCampbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.ORCID 0000-0002-4111-0156
Ramzi HalabiDepartment of Psychiatry, Faculty of Medicine, UT Southwestern Medical Center, Dallas, Texas, USA.ORCID 0000-0002-3784-9556
Kousai GhaouariCampbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.ORCID 0009-0001-6714-6501
Yvonne C Y LauCampbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.ORCID 0009-0007-1399-4541
Martin AldaDepartment of Psychiatry, Dalhousie University, Halifax, Nova Scotia, Canada.ORCID 0000-0001-9544-3944
Arend HintzeDepartment of MicroData Analytics, Dalarna University, Falun, Sweden.
Benoit H MulsantCampbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.ORCID 0000-0002-0303-6450
Abigail OrtizCampbell Family Research Institute, Centre for Addiction and Mental Health, Toronto, Ontario, Canada.ORCID 0000-0001-6886-767X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

BackgroundLarge language models (LLMs) may reduce the burden associated with performing systematic reviews by prescreening abstracts from a literature search for eligibility for inclusion in full-text review.MethodsWe developed an iterative, LLM-based workflow for screening abstracts: after manual specification of eligibility criteria and seed examples, an ensemble of five LLMs deliberates through a Delphi process to classify a batch of abstracts; these labels are used to train a logistic regression model that ranks the remaining abstracts and identifies a new batch of abstracts for LLM escalation until all abstracts are labelled by the LLM or probability thresholds. We tested our workflow on abstracts screened in three published systematic reviews in psychiatry. Our primary endpoint was the recall metric, and secondary endpoint was the work saved over sampling at 95% recall metric (WSS@95%).ResultsIn a dataset on autism biomarkers, 1,655 (35%) of 4,745 retrieved abstracts were judged to be relevant by the original authors. The Delphi-LLM workflow correctly identified 1,605 (97.0%) of these 1,655 abstracts (precision = 54.2%, WSS@95% = 38.1%). The performance metrics were better than non-LLM approaches (recall ≤ 91%, WSS@95 ≤ 26%), and, overall, balanced these metrics optimally compared to single-LLM agents (recall = 84.9-99.9%, WSS@95% = 16.7-39.8%). The recall and work saved metrics were similarly reliable and among the top in two low-prevalence datasets on an attention-deficit hyperactivity disorder treatment review (10% of 2,891 relevant) and a posttraumatic stress disorder trajectory review (7% of 4,453 relevant). For these two datasets, recall was 100.0% and 96.4%, and the WSS@95% was 17.3% and 18.5%, respectively.ConclusionsWe presented the design and validation of a novel abstract screening workflow that centres around a Delphi-style aggregation process to harness the strengths of five open-source LLMs that can be run on consumer-level workstations. This multi-LLM workflow showed acceptable and reliable performance for use as an automated prescreening method to facilitate systematic reviews.

Indexed as

abstract screeningDelphi methodlarge language modelssystematic reviewstext embedding

Identifiers

PMID42059529
PMCPMC13133002

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.