Evidence map›Paper›PMID 41836629›Full record

ArticleDigital health

Performance of latest AI models, RAG, and MCP on lung cancer-related questions.

Xinjie Zhao, Miaomiao Yang, Kang Tian, Hui Jiang, Deyu Guo, Yadong Wang, Jiajun Du

Abstract read
In one paragraph

Article in Digital health. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Xinjie ZhaoInstitute of Oncology, Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China.
Miaomiao YangDepartment of Oncology, Yantai Yuhuangding Hospital Affiliated to Qingdao University, Yantai, China.
Kang TianInstitute of Oncology, Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China.
Hui JiangInstitute of Oncology, Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China.
Deyu GuoInstitute of Oncology, Shandong Provincial Hospital, Shandong University, Jinan, China.
Yadong WangDepartment of Thoracic Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China.
Jiajun DuDepartment of Thoracic Surgery, Shandong Provincial Hospital Affiliated to Shandong First Medical University, Jinan, China.ORCID https://orcid.org/0000-0003-2406-9435

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) have advanced rapidly. However, concerns remain regarding their reliability in clinical settings due to the inherent issues of hallucinations and inadequate referencing. Materials and Methods: We evaluated six current LLMs: GPT-4.1 (GPT), o3, Gemini-2.5-Pro-Preview-0506 (Gemini), Grok-3 (Grok), Qwen3-235B-A22B (Qwen3), and Claude Sonnet 4 (Claude), as well as two technologies that extend LLM capabilities using external knowledge bases: retrieval-augmented generation (RAG) and Model Context Protocol (MCP). Each model was evaluated using 50 questions selected from a 132-question pool developed based on the Chinese Medical Association guideline for clinical diagnosis and treatment of lung cancer (2024 Edition). Three models-Qwen, GPT, and Grok-were further analyzed to assess performance changes with RAG and MCP integration. All responses were independently reviewed by two qualitative evaluators. Results: Overall, o3 achieved the highest accuracy (50%), followed by GPT (48%) and Gemini (48%), then Grok (44%), Qwen (40%), and Claude (36%). However, implementing RAG (LLM-RAG) or MCP (LLM-MCP) significantly improved accuracy, with statistical differences observed between baseline LLMs and their RAG- or MCP-enhanced counterparts. Lexical richness and semantic noise both diminished, whereas the semantic clarity and accuracy of verbs, noun-verb combinations, and content words improved. Conclusions: The six latest LLMs performed similarly on lung cancer-related questions. The integration of RAG or MCP significantly enhanced accuracy while simplifying sentence structure, focusing more on the main topics, and using more accurate vocabulary.

Indexed as

artificial intelligenceChatGPTlarge language modellung cancermodel context protocolretrieval-augmented generation

Identifiers

PMID41836629
PMCPMC12982857

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.