Evidence map›Paper›PMID 42218363›Full record

ArticleBMC medical research methodology2026

Evaluating Elicit's systematic reviews workflow in an umbrella review on air pollution and acute lower respiratory infections: a methodological study for quality appraisal.

Cristina Mazzali, Tiziana Pinciroli, Maria Rosa Valetto, Pietro Dri, Antonio Giampiero Russo, Air and Health Atlas Study Group

Abstract read
In one paragraph

Article in BMC medical research methodology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Cristina MazzaliEpidemiology Unit, Agency for Health Protection (ATS) of Milan, Via Conca del Naviglio, 45, Milan, 20123, Italy. cmazzali@ats-milano.it.ORCID http://orcid.org/0000-0003-2645-9939
Tiziana PinciroliZadig Ltd, Benefit Company, Via Ampère, 59, Milan, 20131, Italy.ORCID http://orcid.org/0009-0006-8581-1284
Maria Rosa ValettoZadig Ltd, Benefit Company, Via Ampère, 59, Milan, 20131, Italy.ORCID http://orcid.org/0000-0002-5195-5208
Pietro DriZadig Ltd, Benefit Company, Via Ampère, 59, Milan, 20131, Italy.ORCID http://orcid.org/0000-0003-3565-4787
Antonio Giampiero RussoEpidemiology Unit, Agency for Health Protection (ATS) of Milan, Via Conca del Naviglio, 45, Milan, 20123, Italy.ORCID http://orcid.org/0000-0002-5681-5861
Air and Health Atlas Study Group

Funding

Ministero della Salute PNC PREVA-2022-12376981
6 · The paper itself

Abstract

backgroundThe exponential growth of scientific publications has increased the complexity of evidence synthesis. Systematic reviews remain essential but highly resource-intensive. Large language models (LLMs) offer new opportunities to support or partially automate key steps of this process. This study evaluates the performance of Elicit's Systematic Reviews workflow in comparison to the traditional methodology, using as reference a published umbrella review on the association between air pollution and acute lower respiratory infections (ALRI).

methodsA parallel workflow was developed to reproduce each phase of the traditional review. Considering article retrieval, for the traditional workflow, articles were retrieved through a Boolean search, while for the AI-assisted workflow a natural-language query submitted to Elicit. Screening was conducted in two steps, emulating the traditional PECOS-based criteria. Full-text evaluation and quality appraisal were performed through Elicit's "data extraction" functionalities. For quality appraisal the validated AMSTAR-2 EH questionnaire was applied.

resultsThe traditional Boolean search identified 324 unique articles. When compared with the 500 records retrieved by Elicit, an overlap of 8% was observed, which prevented a direct, recall-oriented comparison of search performance. To enable a controlled comparison of downstream steps, we applied Elicit's screening and data-extraction functions to the 324 records identified through the Boolean search, using empirically defined screening-score threshold in Elicit to select studies for further evaluation. In the screening on title and abstract, 33 articles were identified through the traditional workflow and 70 through Elicit, 30 articles overlapping (recall 90.9%, precision 42.9%). The full-text screening selected 15 articles with the traditional methodology and 24 with Elicit, all 15 articles from the traditional methodology being included in Elicit selection (recall 100%, precision 62.5%). In the quality assessment, Elicit showed 24.4% disagreement on general items and 30.4% on additional items of the AMSTAR-2 EH. Errors clustered around multi-component questions, items requiring expert interpretation, and information located in supplementary materials.

conclusionsElicit can support several phases of systematic reviews and reduce manual workload, but it cannot independently reproduce the methodological rigor required for high-quality evidence synthesis. At present, LLM-based tools are best positioned as complementary systems within human-supervised workflows.

Indexed as

Air PollutionRespiratory Tract InfectionsSystematic Reviews as TopicWorkflowHumansLarge Language ModelsAMSTAR-2 EHArtificial IntelligenceData extractionEvidence synthesisLarge language modelsSystematic reviews

Identifiers

PMID42218363
PMCPMC13412304

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.