Evidence map›Paper›PMID 40332983›Full record

SynthesisJournal of the American Medical Informatics Association : JAMIA2025

The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review.

Dmitry Scherbakov, Nina Hubig, Vinita Jansari, Alexander Bakumenko, Leslie A Lenert

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of the American Medical Informatics Association : JAMIA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 36 papers, 3 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
36citing papers in PubMed, 3 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

36 citing papers in PubMed, 3 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Article
  5. Article
  6. Review
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Review
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Dmitry ScherbakovBiomedical Informatics Center, Department of Public Health Sciences, Medical University of South Carolina (MUSC), Charleston, SC 29403, United States.ORCID 0009-0005-7274-0934
Nina HubigBiomedical Informatics Center, Department of Public Health Sciences, Medical University of South Carolina (MUSC), Charleston, SC 29403, United States.
Vinita JansariSchool of Computing, Clemson University, Charleston, SC 29634, United States.
Alexander BakumenkoSchool of Computing, Clemson University, Charleston, SC 29634, United States.
Leslie A LenertBiomedical Informatics Center, Department of Public Health Sciences, Medical University of South Carolina (MUSC), Charleston, SC 29403, United States.ORCID 0000-0002-9680-5094

Funding

South Carolina Clinical & Translational Research Institute (SCTR)UL1TR001450 · NCATS · MEDICAL UNIVERSITY OF SOUTH CAROLINA · PI BRADY, KATHLEEN T., FLUME, PATRICK A · 2015 to 2024
$41.1M
SC Biomedical Informatics & Data Science for Health Impact (SC BIDS4Health)T15LM013977 · NLM · MEDICAL UNIVERSITY OF SOUTH CAROLINA · PI Alexander V Alekseyenko, Brian Dean · 2022 to 2026
$1.2M
Biomedical Informatics and Data Science for Health Equity Research SC BIDS4HealthNCATS NIH HHS UL1 TR001450NIH HHS T15 LM013977NIH HHS UL1 TR001450NLM NIH HHS T15 LM013977Smart-state Chair endowment
6 · The paper itself

Abstract

objectivesThis study aims to summarize the usage of large language models (LLMs) in the process of creating a scientific review by looking at the methodological papers that describe the use of LLMs in review automation and the review papers that mention they were made with the support of LLMs. MATERIALS AND

methodsThe search was conducted in June 2024 in PubMed, Scopus, Dimensions, and Google Scholar by human reviewers. Screening and extraction process took place in Covidence with the help of LLM add-on based on the OpenAI GPT-4o model. ChatGPT and Scite.ai were used in cleaning the data, generating the code for figures, and drafting the manuscript.

resultsOf the 3788 articles retrieved, 172 studies were deemed eligible for the final review. ChatGPT and GPT-based LLM emerged as the most dominant architecture for review automation (n = 126, 73.2%). A significant number of review automation projects were found, but only a limited number of papers (n = 26, 15.1%) were actual reviews that acknowledged LLM usage. Most citations focused on the automation of a particular stage of review, such as Searching for publications (n = 60, 34.9%) and Data extraction (n = 54, 31.4%). When comparing the pooled performance of GPT-based and BERT-based models, the former was better in data extraction with a mean precision of 83.0% (SD = 10.4) and a recall of 86.0% (SD = 9.8). DISCUSSION AND

conclusionOur LLM-assisted systematic review revealed a significant number of research projects related to review automation using LLMs. Despite limitations, such as lower accuracy of extraction for numeric data, we anticipate that LLMs will soon change the way scientific reviews are conducted.

Indexed as

Programming LanguagesReview Literature as TopicSystematic Reviews as TopicHumansInformation Storage and RetrievalLarge Language ModelsCovidencelarge language modelsreview automationscoping reviewsystematic review

Identifiers

PMID40332983
PMCPMC12089777

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.