Evidence map›Paper›PMID 42166792›Full record

ArticleJournal of medical Internet research2026

Benchmarking Large Language Models and Prompt Engineering Strategies in Microsatellite Instability Cancers: Evaluation Study.

Yuxin Zhang, Jie Song, Cheng Bi, Xin Zheng, Zhichuan Xu, Dan Cao, Bairong Shen

Abstract readEvaluation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Yuxin Zhang *Department of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0009-0000-3067-6341
Jie Song *Department of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0009-0007-1606-0605
Cheng BiDepartment of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0000-0002-8076-2704
Xin ZhengDepartment of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0009-0000-9300-8949
Zhichuan XuDepartment of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0009-0002-2690-2961
Dan CaoDepartment of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0000-0003-2755-7258
Bairong ShenDepartment of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.ORCID http://orcid.org/0000-0003-2899-1531

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The reliability of general-purpose large language models (LLMs) for complex clinical tasks in specialized domains such as microsatellite instability (MSI) cancers remains critically uncharacterized. The absence of a domain-specific benchmark to evaluate and guide the optimization of their capabilities across diverse clinical tasks poses unevaluated risks to patient safety. Objective: This study aimed to develop and validate Microsatellite Instability Cancer Benchmark (MSIC-Bench), a novel, two-tiered benchmark for MSI cancer, covering both consensus and frontier knowledge. Using this framework, we aimed to systematically assess LLM performance across various prompting strategies, identify task-specific weaknesses, and reveal effective pathways for performance improvement. Methods: We developed MSIC-Bench, a 511-question benchmark derived from clinical guidelines and a curated knowledge base. Three state-of-the-art LLMs (GPT-4o, [OpenAI], Gemini 2.5 Pro [Google], and Claude Opus 4 [Anthropic]) were evaluated using 4 prompting strategies, including vanilla, chain-of-thought, reflection of thoughts, and retrieval-augmented generation (RAG), under both multiple-choice and open-ended modalities. Performance was assessed on accuracy, safety (honesty), error composition, and token usage. Results: LLMs demonstrated a significant "scaffolding effect," with accuracy dropping substantially in open-ended scenarios. For non-RAG strategies, the primary failure mode was an internal knowledge deficit. The integration of RAG proved to be the most effective intervention. A domain-aligned RAG strategy not only significantly improved accuracy in complex decision-making tasks but also fundamentally shifted the system's primary bottleneck from knowledge deficits to retrieval failures. In terms of safety, RAG induced a favorable shift from high-risk fabrications to safer refusals, though this introduced a safety-utility trade-off in the form of false refusals. Notably, our hybrid-RAG configuration, which combined both knowledge sources, demonstrated the most robust and generalizable performance across all tasks. Conclusions: Current LLMs lack the specialized knowledge required for reliable application in MSI oncology. A well-designed RAG architecture is the pivotal intervention to address this gap. However, its success is not automatic; it transforms the nature of system failure, making retrieval precision and knowledge base quality the new critical determinants of performance and safety. Our findings establish a clear directive for developing trustworthy clinical artificial intelligence: focus must shift toward optimizing the retrieval component and curating high-quality, comprehensive knowledge sources. MSIC-Bench provides a robust framework to guide these future efforts.

Indexed as

BenchmarkingLarge Language ModelsMicrosatellite InstabilityNeoplasmsHumansbenchmarkcancerlarge language modelLLMmicrosatellite instabilityprompt engineering

Identifiers

PMID42166792
PMCPMC13193672

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.