Evidence map›Paper›PMID 38879534›Full record

ArticleSystematic reviews2024

Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain.

Fabio Dennstädt, Johannes Zink, Paul Martin Putora, Janna Hastings, Nikola Cihoric

Abstract read
In one paragraph

Article in Systematic reviews, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 56 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
56citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

56 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. In Reply to Sengul I and Sengul D.Advances in radiation oncology · 2026
    Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. CLEAR: Comparative Letter Examination and Analysis for Red Flags.Journal of graduate medical education · 2026
    Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Fabio DennstädtDepartment of Radiation Oncology, Cantonal Hospital of St. Gallen, St. Gallen, Switzerland. fabiodennstaedt@gmx.de.ORCID 0000-0002-5374-8720
Johannes ZinkInstitute for Computer Science, University of Würzburg, Würzburg, Germany.
Paul Martin PutoraDepartment of Radiation Oncology, Cantonal Hospital of St. Gallen, St. Gallen, Switzerland.
Janna HastingsInstitute for Implementation Science in Health Care, University of Zurich, Zurich, Switzerland.
Nikola CihoricDepartment of Radiation Oncology, Inselspital, Bern University Hospital and University of Bern, Bern, Switzerland.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundSystematically screening published literature to determine the relevant publications to synthesize in a review is a time-consuming and difficult task. Large language models (LLMs) are an emerging technology with promising capabilities for the automation of language-related tasks that may be useful for such a purpose.

methodsLLMs were used as part of an automated system to evaluate the relevance of publications to a certain topic based on defined criteria and based on the title and abstract of each publication. A Python script was created to generate structured prompts consisting of text strings for instruction, title, abstract, and relevant criteria to be provided to an LLM. The relevance of a publication was evaluated by the LLM on a Likert scale (low relevance to high relevance). By specifying a threshold, different classifiers for inclusion/exclusion of publications could then be defined. The approach was used with four different openly available LLMs on ten published data sets of biomedical literature reviews and on a newly human-created data set for a hypothetical new systematic literature review.

resultsThe performance of the classifiers varied depending on the LLM being used and on the data set analyzed. Regarding sensitivity/specificity, the classifiers yielded 94.48%/31.78% for the FlanT5 model, 97.58%/19.12% for the OpenHermes-NeuralChat model, 81.93%/75.19% for the Mixtral model and 97.58%/38.34% for the Platypus 2 model on the ten published data sets. The same classifiers yielded 100% sensitivity at a specificity of 12.58%, 4.54%, 62.47%, and 24.74% on the newly created data set. Changing the standard settings of the approach (minor adaption of instruction prompt and/or changing the range of the Likert scale from 1-5 to 1-10) had a considerable impact on the performance.

conclusionsLLMs can be used to evaluate the relevance of scientific publications to a certain review topic and classifiers based on such an approach show some promising results. To date, little is known about how well such systems would perform if used prospectively when conducting systematic literature reviews and what further implications this might have. However, it is likely that in the future researchers will increasingly use LLMs for evaluating and classifying scientific publications.

Indexed as

Natural Language ProcessingBiomedical ResearchLanguageSystematic Reviews as TopicBiomedicineLarge language modelsNatural language processingSystematic literature reviewTitle and abstract screening

Identifiers

PMID38879534
PMCPMC11180407

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.