Evidence map›Paper›PMID 40567772›Full record

ArticlePeerJ. Computer science2025

Validation of automated paper screening for esophagectomy systematic review using large language models.

Rashi Ramchandani, Eddie Guo, Esra Rakab, Jharna Rathod, Jamie Strain, William Klement, Risa Shorr, Erin Williams, Daniel Jones, Sebastien Gilbert

Abstract read
In one paragraph

Article in PeerJ. Computer science, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Review
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Rashi RamchandaniDepartment of Medicine, University of Ottawa, Ottawa, Ontario, Canada.
Eddie GuoCumming School of Medicine, University of Calgary, Calgary, Alberta, Canada.
Esra RakabDepartment of Medicine, University of Ottawa, Ottawa, Ontario, Canada.
Jharna RathodDepartment of Medicine, University of Ottawa, Ottawa, Ontario, Canada.
Jamie StrainOttawa Hospital Research Institute, Ottawa, Ontario, Canada.ORCID 0000-0002-4023-5715
William KlementOttawa Hospital Research Institute, Ottawa, Ontario, Canada.
Risa ShorrLibrary and Learning Services, The Ottawa Hospital, Ottawa, Ontario, Canada.
Erin WilliamsDivision of General Surgery, Department of Surgery, The Ottawa Hospital, Ottawa, Ontario, Canada.ORCID 0000-0001-5418-814X
Daniel JonesDivision of General Surgery, Department of Surgery, The Ottawa Hospital, Ottawa, Ontario, Canada.
Sebastien GilbertDivision of General Surgery, Department of Surgery, The Ottawa Hospital, Ottawa, Ontario, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) offer a potential solution to the labor-intensive nature of systematic reviews. This study evaluated the ability of the GPT model to identify articles that discuss perioperative risk factors for esophagectomy complications. To test the performance of the model, we tested GPT-4 on narrower inclusion criterion and by assessing its ability to discriminate relevant articles that solely identified preoperative risk factors for esophagectomy. Methods: A literature search was run by a trained librarian to identify studies ( Results: The agreement between the GPT model and human decision was 85.58% for perioperative factors and 78.75% for preoperative factors. The AUC value was 0.87 and 0.75 for the perioperative and preoperative risk factors query, respectively. In the evaluation of perioperative risk factors, the GPT model demonstrated a high recall for included studies at 89%, a positive predictive value of 74%, and a negative predictive value of 84%, with a low false positive rate of 6% and a macro-F1 score of 0.81. For preoperative risk factors, the model showed a recall of 67% for included studies, a positive predictive value of 65%, and a negative predictive value of 85%, with a false positive rate of 15% and a macro-F1 score of 0.66. The interobserver reliability was substantial, with a kappa score of 0.69 for perioperative factors and 0.61 for preoperative factors. Despite lower accuracy under more stringent criteria, the GPT model proved valuable in streamlining the systematic review workflow. Preliminary evaluation of inclusion and exclusion justification provided by the GPT model were reported to have been useful by study screeners, especially in resolving discrepancies during title and abstract screening. Conclusion: This study demonstrates promising use of LLMs to streamline the workflow of systematic reviews. The integration of LLMs in systematic reviews could lead to significant time and cost savings, however caution must be taken for reviews involving stringent a narrower and exclusion criterion. Future research is needed and should explore integrating LLMs in other steps of the systematic review, such as full text screening or data extraction, and compare different LLMs for their effectiveness in various types of systematic reviews.

Indexed as

Abstract screeningChatGPTLarge language modelScreeningSystematic review

Identifiers

PMID40567772
PMCPMC12190591

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.