Evidence map›Paper›PMID 41127326›Full record

ArticleCochrane evidence synthesis and methods2025

Enhancing Evidence Synthesis Efficiency: Leveraging Large Language Models and Agentic Workflows for Optimized Literature Screening.

Bing Hu, Emmalie Tomini, Tricia Corrin, Kusala Pussegoda, Elias Sandner, Andre Henriques, Alice Simniceanu, Luca Fontana, Andreas Wagner, Stephanie Brazeau and 1 more

Abstract read
In one paragraph

Article in Cochrane evidence synthesis and methods, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Bing HuData Management, Innovation and Analytics, Data, Surveillance and Foresight, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0009-0006-2913-6659
Emmalie TominiData Management, Innovation and Analytics, Data, Surveillance and Foresight, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0009-0009-2809-1075
Tricia CorrinNational Microbiology Laboratory Public Health Risk Sciences, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0000-0001-5514-1589
Kusala PussegodaNational Microbiology Laboratory Public Health Risk Sciences, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0009-0007-8383-827X
Elias SandnerTechnical University of Graz Graz Austria.ORCID https://orcid.org/0009-0007-9855-4923
Andre HenriquesCERN (European Organization for Nuclear Research) Graz Austria.ORCID https://orcid.org/0000-0003-1521-3423
Alice SimniceanuHealth Emergencies Programme WHE, World Health Organization Geneva Switzerland.ORCID https://orcid.org/0000-0003-4068-6177
Luca FontanaHealth Emergencies Programme WHE, World Health Organization Geneva Switzerland.ORCID https://orcid.org/0000-0002-8614-4114
Andreas WagnerCERN (European Organization for Nuclear Research) Graz Austria.ORCID https://orcid.org/0000-0001-9589-2635
Stephanie BrazeauData Management, Innovation and Analytics, Data, Surveillance and Foresight, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0009-0004-6190-2160
Lisa WaddellNational Microbiology Laboratory Public Health Risk Sciences, Public Health Agency of Canada Ottawa Ontario Canada.ORCID https://orcid.org/0000-0003-4887-5124

Funding

World Health Organization 001
6 · The paper itself

Abstract

Background: Public health events of international concern highlight the need for up-to-date evidence curated using sustainable processes that are accessible. In development of the Global Repository of Epidemiological Parameters (grEPI) we explore the performance of an agentic-AI assisted pipeline (GREP-Agent) for screening evidence which capitalizes on recent advancements in large language models (LLMs). Methods: In this study, the performance of the GREP-Agent was evaluated on a data set of 2000 citations from a systematic review on measles using four LLMs (GPT4o, GPT4o-mini, Llama3.1, and Phi4). The GREP-Agent framework integrates multiple LLMs and human feedback to fine-tune its performance, optimize workload reduction and accuracy in screening research articles. The impact on performance of each part of this Agentic-AI system is presented and measured by accuracy, precision, recall, and F1-score metrics. Results: The results show how each phase of the GREP-Agent system incrementally improves accuracy regardless of the LLM. We found that GREP-Agent was able to increase sensitivity across a broad range of open source and proprietary LLMs to 84.2%-88.9% after fine-tuning and to 86.4%-95.3% by varying workload reduction strategies. Performance was significantly impacted by the clarity of the screening questions and setting thresholds for optimized workload reduction strategies. Conclusions: The GREP-Agent shows promise in improving the efficiency and effectiveness of evidence synthesis in dynamic public health contexts. Further development and refinement of adaptable human-in-the-loop AI systems for screening literature are essential to support future public health response activities, while maintaining a human-centric approach.

Identifiers

PMID41127326
PMCPMC12538819

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.