Evidence map›Paper›PMID 41773687›Full record

ArticleJMIR cancer2026

ChatGPT Versus DeepSeek for Breast Cancer Information Retrieval: Quantitative Comparative Study.

Rima Hajjo, Dima A Sabbah, Sanaa K Bardaweel

Abstract readComparative Study
In one paragraph

Article in JMIR cancer, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Rima HajjoDepartment of Pharmacy, Faculty of Pharmacy, Al-Zaytoonah University of Jordan, Airport Street, P.O. Box 130, Amman, 11733, Jordan, 962 64291511.ORCID 0000-0002-7090-5425
Dima A SabbahDepartment of Pharmacy, Faculty of Pharmacy, Al-Zaytoonah University of Jordan, Airport Street, P.O. Box 130, Amman, 11733, Jordan, 962 64291511.ORCID 0000-0003-1428-5097
Sanaa K BardaweelDepartment of Pharmaceutical Sciences, School of Pharmacy, University of Jordan, Amman, Jordan.ORCID 0000-0002-4823-0708

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Artificial intelligence (AI) is increasingly used to generate medical content, yet its performance in delivering clinically relevant and reliable information remains underexplored, especially in complex areas such as breast cancer. Objective: This study aimed to compare ChatGPT-4.0 and DeepSeek-V3 in generating breast cancer information, focusing on readability, content quality, and citation reliability. Methods: On the basis of publicly available patient education materials, 10 frequently asked questions were selected. Each model generated 60 responses. Three expert reviewers rated each response using a 7-point Likert scale across 5 dimensions (ie, accuracy, completeness, clarity, depth and insight, and alignment with expert answers). Readability was assessed using Flesch-Kincaid Grade Level scores. Information reliability was evaluated through interrater agreement metrics, including Cohen κ and Fleiss κ. Paired t tests were used for statistical comparisons. Results: AI models produced significantly more readable content than expert references (mean Flesch-Kincaid Grade Level difference -2.60; P<.001). ChatGPT-4.0 responses were more stylistically consistent with a median Flesch-Kincaid Grade Level score of 10.66 (IQR 0.98), whereas DeepSeek-V3 showed greater variability with a median Flesch-Kincaid Grade Level score of 10.17 (IQR 1.41). Content quality scores were DeepSeek-V3 achieving a higher mean score than ChatGPT-4.0 (6.22 [SD 0.43] vs 6.01 [SD 0.49]). In the multiresponse analysis, DeepSeek-V3 demonstrated a statistically significant advantage in accuracy (P=.041), while differences across other criteria were not statistically significant (P>.05). Human raters showed almost perfect agreement when judging source reliability (Fleiss κ=0.842 for ChatGPT's citations and 0.935 for DeepSeek's citations). Agreement between each model's citation reliability scores and the expert majority was substantial for ChatGPT (Cohen κ=0.665) and higher for DeepSeek (Cohen κ=0.782). Conclusions: Both models generated readable and clinically relevant content with comparable overall performance. ChatGPT provided more consistent readability, while DeepSeek offered more diverse references with stronger alignment to expert ratings. Continued evaluation and quality assurance are essential for the responsible clinical use of AI-generated content.

Indexed as

Breast NeoplasmsInformation Storage and RetrievalArtificial IntelligenceComprehensionFemaleGenerative Artificial IntelligenceHumansLarge Language ModelsReproducibility of ResultsAIartificial intelligencebreast cancerChatGPTDeepSeeklarge language modelsLLM

Identifiers

PMID41773687
PMCPMC12954694

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.