Evidence map›Paper›PMID 41094286›Full record

ArticleAnnals of surgical oncology2026

Evaluating the Clinical Competence of Large Language Models in Prostate Cancer Management: A Comparative Study of DeepSeek-R1 and ChatGPT.

Rongkang Li, Anguo Zhao, Lei Peng, Hongjin Shi, Jiadong Zhao, Zhilin Li, Rui Liang, Haifeng Wang

Abstract readComparative Study
PubMed Publisher
In one paragraph

Article in Annals of surgical oncology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Review
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Rongkang LiDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China.
Anguo ZhaoDepartment of Urology, The Fourth Affiliated Hospital of Soochow University, Medical Center of Soochow University, Suzhou Dushu Lake Hospital, Suzhou, People's Republic of China.
Lei PengDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China.
Hongjin ShiDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China.
Jiadong ZhaoDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China.
Zhilin LiDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China.
Rui LiangDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China. 18326900957@163.com.
Haifeng WangDepartment of Urology, The Second Affiliated Hospital of Kunming Medical University, Yunnan Institute of Urology, Kunming, Yunnan, People's Republic of China. wanghaifeng@kmmu.edu.cn.

Funding

Scientific Research Fund Project of Education Department of Yunnan Province No. 2024Y231
6 · The paper itself

Abstract

backgroundLarge language models (LLMs) have gained prominence in medical applications, yet their performance in specialized clinical tasks remains underexplored. Prostate cancer, a complex malignancy requiring guideline-based management, presents a rigorous testbed for evaluating artificial intelligence (AI)-assisted decision-making. This study compared the clinical accuracy, reasoning ability, and language quality of DeepSeek-R1 and ChatGPT variants in addressing prostate cancer diagnosis and treatment.

methodsA dataset of 98 prostate cancer multiple-choice questions from MedQA, MedMCQA, and China's National Medical Licensing Examination was constructed, alongside three real-world clinical cases. Responses were generated by five LLMs (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, -o3, -o4-mini) and evaluated for accuracy across three repeated runs. For case-based simulations, only R1 and o3 were compared with practicing urologists. A Clinical Decision Quality Assessment Scale (CDQAS) assessed outputs across four domains: readability, medical knowledge accuracy, diagnostic test appropriateness, and logical coherence. Blinded scoring was performed by senior urologic oncologists. Statistical analyses used one-way ANOVA with GraphPad Prism v10.1.2, Boston, Massachusetts, USA.

resultsDeepSeek-R1 achieved the highest accuracy (96.60 %) on multiple-choice tasks, significantly outperforming the other models (p < 0.05 to <0.0001). In simulated case evaluations, both R1 and o3 performed comparably with physicians in overall readability and diagnostic appropriateness. Whereas R1 demonstrated superior guideline compliance and evidence-based reasoning, o3 showed advantages in workflow clarity, sequencing, and response fluency. However, o3 generated fewer explicit errors than R1. Human clinicians maintained strengths in terminology precision and logical reasoning.

conclusionDeepSeek-R1 and ChatGPT-o3 exhibit complementary strengths in prostate cancer clinical decision-making, with R1 favoring factual accuracy and o3 excelling in expressive clarity. Although both models approach human-level performance in structured evaluations, human oversight and continued domain-specific optimization remain essential for their safe and effective integration into clinical workflows.

Indexed as

Artificial IntelligenceClinical CompetenceClinical Decision-MakingLanguageProstatic NeoplasmsGenerative Artificial IntelligenceHumansLarge Language ModelsMaleChatGPTClinical decision-makingDeepSeek-R1Large language models (LLMs)Prostate cancer

Identifiers

PMID41094286

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.