Evidence map›Paper›PMID 42166797›Full record

SynthesisJournal of medical Internet research2026

Large Language Models in Colorectal Cancer Care and Clinical Decision Support: Systematic Review.

Jinglei Tian, Qifeng Lou, Xue Wang, Hangying Xu, Huiting Mei, Yanli Yu

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Jinglei TianZhejiang Chinese Medical University, Zhejiang Chinese Medical University, Hangzhou, China, Hangzhou, Zhejiang, China.ORCID http://orcid.org/0009-0009-3104-7606
Qifeng LouDepartment of Gastroenterology, Hangzhou First People's Hospital, Hangzhou, China, Hangzhou, Zhejiang, China, 86 15267498545.ORCID http://orcid.org/0000-0002-8303-8340
Xue WangDepartment of Gastroenterology, Hangzhou First People's Hospital, Hangzhou, China, Hangzhou, Zhejiang, China, 86 15267498545.ORCID http://orcid.org/0009-0006-6936-7224
Hangying XuZhejiang Chinese Medical University, Zhejiang Chinese Medical University, Hangzhou, China, Hangzhou, Zhejiang, China.ORCID http://orcid.org/0009-0002-6611-7220
Huiting MeiZhejiang Chinese Medical University, Zhejiang Chinese Medical University, Hangzhou, China, Hangzhou, Zhejiang, China.ORCID http://orcid.org/0009-0007-7292-7234
Yanli YuZhejiang Chinese Medical University, Zhejiang Chinese Medical University, Hangzhou, China, Hangzhou, Zhejiang, China.ORCID http://orcid.org/0009-0001-3004-515X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Colorectal cancer (CRC) is a leading cause of cancer morbidity and mortality worldwide. The complexity of guideline-concordant care and unstructured clinical data has driven demand for decision-support tools. Large language models (LLMs) show promise for processing clinical data and patient-provider communication, yet evidence is fragmented, and a CRC-specific synthesis across the full care continuum is lacking. Objective: This systematic review evaluates the current applications, performance determinants, and clinical implications of LLMs across the continuum of CRC care. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), we searched 6 databases (PubMed, Embase, Web of Science, Scopus, CINAHL, Cochrane) through April 1, 2026. Eligible studies were peer-reviewed original investigations of LLMs on CRC tasks with extractable outcomes; reviews, editorials, and abstracts were excluded. Two reviewers assessed quality with QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2), PROBAST (prediction model risk of bias assessment tool), and ROBINS-I (Risk of Bias in Nonrandomized Studies - of Interventions). Data on model types, applications, prompts, input/output formats, and outcomes were analyzed descriptively, with narrative synthesis per synthesis without meta-analysis (SWiM) guidelines. Results: Of 8880 records, 37 studies met inclusion criteria (2023-2026), mostly from China and the United States, with GPT series most frequently evaluated. Overall risk of bias was low in 10/37 studies (27.0%), moderate in 14/37 (37.8%), unclear in 7/37 (18.9%), and high or serious in 6/37 (16.2%). Problematic domains included outcome measurement, intervention classification, patient selection, and lack of blinded assessment. LLMs showed utility in automating data extraction from clinical texts, supporting patient education, aiding diagnosis, and assisting clinical decision-making, with emerging visual interpretation and multimodal capacities. Domain-specific and multimodal models showed advantages over general-purpose models in certain tasks. Performance was significantly influenced by prompt design, from zero-shot queries to fine-tuning. Despite efficiency and outcome benefits, challenges persist regarding methodological quality, data privacy, and generalizability. Conclusions: This review provides an integrative framework synthesizing evidence across study designs and LLM categories in CRC care. Unlike prior reviews addressing gastroenterology broadly or limited to one design, it covers the full CRC continuum and, for the first time, comparatively evaluates general-purpose, domain-specific, and multimodal LLMs, clarifying how prompt engineering and heterogeneous metrics shape outcomes. Although findings support LLMs' clinical potential, results must be interpreted cautiously, given low overall evidence quality. Most studies lacked safeguards against bias-blinded assessment, confounder adjustment, or prospective multicenter validation. Substantial heterogeneity across tasks, LLM types, prompts, reference standards, and outcomes means reported advantages cannot be generalized. Future work should prioritize real-world integration via prospective multicenter validation, robust privacy frameworks, and rigorous human oversight. Amid rising global CRC burden and health care disparities, this review informs clinical translation, equitable scaling, and policy on LLM deployment.

Indexed as

Colorectal NeoplasmsDecision Support Systems, ClinicalLarge Language ModelsHumansartificial intelligencecolorectal cancergastroenterologylarge language modelsPRISMAsystematic review

Identifiers

PMID42166797
PMCPMC13193707

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.