ArticleSystematic reviews2024
Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain.
Article in Systematic reviews, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 56 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
56 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Artificial Intelligence for Evidence Synthesis of Emerging Biologics to Improve Skeletal Health in Osteogenesis Imperfecta: Systematic Review and Meta-Analysis.Journal of medical Internet research · 2026Pooled it
- A comparative study of screening performance between abstrackr and GPT models: Systematic review and contextual analysis.BMC medical informatics and decision making · 2025Pooled it
- Ensemble of transformers for depression emotion classification.Cognitive neurodynamics · 2026Article
- Article
- Introduction to Concepts in Artificial Intelligence and Machine Learning for Pharmacoepidemiologists: Large Language Models.Pharmacoepidemiology and drug safety · 2026Article
- Leveraging prompt-driven generative AI for systematic reviews in digital psychiatry: A stage-matched comparative proof-of-concept for healthcare researchers and clinicians.PLOS digital health · 2026Article
- Artificial Intelligence Resources for the Screening of Titles and Abstracts in Systematic Reviews: A Scoping Review.Cochrane evidence synthesis and methods · 2026Review
- Toward Automating the Selection of Articles Reporting EQ-5D Data for Systematic Literature Reviews Using Large Language Models: Algorithm Development and Evaluation Study.JMIR formative research · 2026Article
- Leveraging generative artificial intelligence for the development of non-interventional research study protocols: a proof-of-concept feasibility study.BMC medical research methodology · 2026Article
- In Reply to Sengul I and Sengul D.Advances in radiation oncology · 2026Article
- Implementing a Resource-Light and Low-Code Large Language Model System for Information Extraction from Mammography Reports: A Pilot Study.Journal of imaging informatics in medicine · 2026Article
- Automated full-text screening and accelerated reviews using large language models with context-aware agents: an exploratory analysis in biomarker research.European heart journal. Digital health · 2026Article
- Evaluating Elicit's systematic reviews workflow in an umbrella review on air pollution and acute lower respiratory infections: a methodological study for quality appraisal.BMC medical research methodology · 2026Article
- To include or not to include? A prescription from the pharmacy on how to use active learning-assisted screening in systematic reviews.Systematic reviews · 2026Article
- Optimizing document retrieval using massive text embeddings and LLM prompt engineering.Systematic reviews · 2026Article
- Automatically detecting trends and open questions from mental health publications: a Wellcome-funded GALENOS project.BMJ mental health · 2026Article
- CLEAR: Comparative Letter Examination and Analysis for Red Flags.Journal of graduate medical education · 2026Article
- Beyond human gold standards: A multimodel framework for automated abstract classification and information extraction.Research synthesis methods · 2026Article
- Compact large language models for title and abstract screening in systematic reviews: An assessment of feasibility, accuracy, and workload reduction.Research synthesis methods · 2026Article
- The landscape of artificial intelligence tools and platforms for evidence synthesis: a scoping review.Systematic reviews · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundSystematically screening published literature to determine the relevant publications to synthesize in a review is a time-consuming and difficult task. Large language models (LLMs) are an emerging technology with promising capabilities for the automation of language-related tasks that may be useful for such a purpose.
methodsLLMs were used as part of an automated system to evaluate the relevance of publications to a certain topic based on defined criteria and based on the title and abstract of each publication. A Python script was created to generate structured prompts consisting of text strings for instruction, title, abstract, and relevant criteria to be provided to an LLM. The relevance of a publication was evaluated by the LLM on a Likert scale (low relevance to high relevance). By specifying a threshold, different classifiers for inclusion/exclusion of publications could then be defined. The approach was used with four different openly available LLMs on ten published data sets of biomedical literature reviews and on a newly human-created data set for a hypothetical new systematic literature review.
resultsThe performance of the classifiers varied depending on the LLM being used and on the data set analyzed. Regarding sensitivity/specificity, the classifiers yielded 94.48%/31.78% for the FlanT5 model, 97.58%/19.12% for the OpenHermes-NeuralChat model, 81.93%/75.19% for the Mixtral model and 97.58%/38.34% for the Platypus 2 model on the ten published data sets. The same classifiers yielded 100% sensitivity at a specificity of 12.58%, 4.54%, 62.47%, and 24.74% on the newly created data set. Changing the standard settings of the approach (minor adaption of instruction prompt and/or changing the range of the Likert scale from 1-5 to 1-10) had a considerable impact on the performance.
conclusionsLLMs can be used to evaluate the relevance of scientific publications to a certain review topic and classifiers based on such an approach show some promising results. To date, little is known about how well such systems would perform if used prospectively when conducting systematic literature reviews and what further implications this might have. However, it is likely that in the future researchers will increasingly use LLMs for evaluating and classifying scientific publications.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.