Evidence map›Paper›PMID 42346209›Full record

ArticleCurrent oncology (Toronto, Ont.)2026

Assessment of Large Language Models in Colorectal Cancer Multidisciplinary Tumor Board Decision-Making: A Retrospective Single-Center Comparison of Guideline-Integrated General-Purpose vs. Domain-Specialized Models.

Aydan Farzaliyeva, Mehmet Nezir Ramazanoglu, Arzu Oguz, Ozden Altundag, Zafer Akcali

Abstract readComparative Study
In one paragraph

Article in Current oncology (Toronto, Ont.), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Aydan FarzaliyevaDivision of Medical Oncology, Department of Internal Medicine, Faculty of Medicine, Baskent University, 06490 Ankara, Türkiye.ORCID 0000-0001-8504-4798
Mehmet Nezir RamazanogluDivision of Medical Oncology, Department of Internal Medicine, Faculty of Medicine, Baskent University, 06490 Ankara, Türkiye.ORCID 0009-0007-4976-182X
Arzu OguzDivision of Medical Oncology, Department of Internal Medicine, Faculty of Medicine, Baskent University, 06490 Ankara, Türkiye.
Ozden AltundagDivision of Medical Oncology, Department of Internal Medicine, Faculty of Medicine, Baskent University, 06490 Ankara, Türkiye.
Zafer AkcaliDivision of Medical Oncology, Department of Internal Medicine, Faculty of Medicine, Baskent University, 06490 Ankara, Türkiye.ORCID 0000-0003-2473-4431

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs) are emerging as clinical decision-support tools in oncology, yet their ability to generate reliable treatment recommendations in real-world multidisciplinary tumor board (MTB) settings remains uncertain, particularly for complex colorectal cancer (CRC).

methodsIn this retrospective study, 300 consecutive adult CRC cases discussed at a tertiary MTB were evaluated. Standardized de-identified case summaries were independently submitted to Gemini 2.5 (general-purpose, guideline-integrated) and MedGemma 27B (domain-specialized; T = 0.0 and T = 1.0). Concordance with MTB decisions was assessed using weighted Cohen's kappa (κ), accuracy, F1 score, and recall. Safety was adjudicated by blinded senior MTB members using a three-tier risk framework. Importantly, Gemini 2.5 was evaluated in a guideline-integrated setting, whereas MedGemma operated without external guideline retrieval, introducing a predefined asymmetry in knowledge augmentation.

resultsGemini 2.5 demonstrated substantial agreement (κ = 0.792,

conclusionsA guideline-integrated general-purpose LLM demonstrated superior concordance and safety compared with a domain-specialized model operating without external retrieval, supporting its adjunctive use within MTBs while preserving expert clinical judgment.

Indexed as

Clinical Decision-MakingColorectal NeoplasmsDecision Support Systems, ClinicalFemaleHumansLarge Language ModelsMaleMiddle AgedRetrospective Studiesartificial intelligenceclinicalcolorectal neoplasmsdecision support systemsgeminilarge language modelsMedGemmamultidisciplinary tumor board

Identifiers

PMID42346209
PMCPMC13297846

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.