Evidence map›Paper›PMID 41821914›Full record

ArticleCentral European journal of urology2026

Evaluation of DeepSeek-R1 and ChatGPT-4o as educational sources for upper tract urothelial carcinoma.

Wojciech Krajewski, Jan Łaszkiewicz, Łukasz Biesiadecki, Wojciech Tomczak, Łukasz Nowak, Piotr Łaszkiewicz, Joanna Chorbińska, Francesco Del Giudice, Benjamin I Chung, Tomasz Szydełko

Abstract read
In one paragraph

Article in Central European journal of urology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Wojciech KrajewskiDepartment of Minimally Invasive and Robotic Urology, University Center of Excellence in Urology, Wroclaw Medical University, Poland.
Jan ŁaszkiewiczUniversity Centre of Excellence in Urology, Wroclaw Medical University, Poland.
Łukasz BiesiadeckiUniversity Centre of Excellence in Urology, Wroclaw Medical University, Poland.
Wojciech TomczakUniversity Centre of Excellence in Urology, Wroclaw Medical University, Poland.
Łukasz NowakDepartment of Minimally Invasive and Robotic Urology, University Center of Excellence in Urology, Wroclaw Medical University, Poland.
Piotr ŁaszkiewiczNova School of Science and Technology, Universidade Nova de Lisboa, Lisbon, Portugal.
Joanna ChorbińskaDepartment of Minimally Invasive and Robotic Urology, University Center of Excellence in Urology, Wroclaw Medical University, Poland.
Francesco Del GiudiceDepartment of Maternal Infant and Urologic Sciences, "Sapienza" University of Rome, Policlinico Umberto I Hospital, Rome, Italy.
Benjamin I ChungDepartment of Urology, Stanford University School of Medicine, Stanford, CA, United States.
Tomasz SzydełkoUniversity Centre of Excellence in Urology, Wroclaw Medical University, Poland.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Introduction: Upper tract urothelial carcinoma (UTUC) is associated with poor survival outcomes. Therefore, providing reliable information about UTUC is crucial. Recently, chatbots powered by large language models have become a widely used information source. Our aim was to evaluate and compare responses generated by ChatGPT-4o and DeepSeek-R1 to patient-important questions regarding UTUC. Material and methods: A set of 43 questions assigned into four categories (general information, symptoms and diagnosis, treatment, prognosis) was curated. Each question was entered into DeepSeek-R1 and ChatGPT-4o. Answers were rated by two urologists using a scale from 1 (completely incorrect) to 4 (fully correct). The median score was calculated for each question. Median scores ≥3 were considered accurate. The repeatability of responses was evaluated using cosine similarity. The number of words in responses was counted. Results: The median scores for DeepSeek-R1 and ChatGPT-4o were both 3.5. There was no statistically significant difference between the scores assigned to two chatbots for all questions (p = 0.35), nor for any particular category.DeepSeek-R1 and ChatGPT-4o provided satisfactory answers for 93% and 91% of the evaluated questions, respectively. No potentially dangerous information was found. Both models consistently generated responses with moderate-high similarity (cosine similarity >0.5), except in one query. Finally, DeepSeek-R1 provided significantly longer answers than ChatGPT-4o (p <0.001). Conclusions: Both DeepSeek-R1 and ChatGPT-4o predominantly provide satisfactory responses to patient-important questions about UTUC. Artificial intelligence chatbots demonstrate potential as the first-line information sources for patients but struggle with highly specialized inquiries and thus cannot replace expert medical advice.

Indexed as

AIartificial intelligenceChatGPTDeepSeekupper tract urothelial carcinoma

Identifiers

PMID41821914
PMCPMC12976754

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-SA
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.