Evidence map›Paper›PMID 40849657›Full record

ArticleBMC oral health2025

Evaluation of deepseek, gemini, ChatGPT-4o, and perplexity in responding to salivary gland cancer.

Ahmed Bashah, Abdulkhaleq Salem, Ali Al-Waqeerah, Eslam Ghaleb, Natheer Wahan, Ahmed Awad, Omran Al-Tos, Gang Chen

Abstract readEvaluation Study
In one paragraph

Article in BMC oral health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. AI in respiratory care: findings from the GOLD report.Journal of translational medicine · 2026
    Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Ahmed BashahDepartment of Stomatology, The First Affiliated Hospital of Dalian Medical University, No.222, Zhongshan Road, Dalian, 116011, Liaoning, PR China.
Abdulkhaleq SalemDepartment of Pharmacology, Dalian Medical University, Dalian, 116011, Liaoning, China.
Ali Al-WaqeerahDepartment of Respiratory, The First Affiliated Hospital of Dalian Medical University, Dalian, 116011, Liaoning, China.
Eslam GhalebDepartment of Biochemistry and Molecular Biology, Dalian Medical University, Dalian, 116011, Liaoning, China.
Natheer WahanDepartment of Pharmacology, Dalian Medical University, Dalian, 116011, Liaoning, China.
Ahmed AwadDepartment of Stomatology, Dalian Medical University, Dalian, 116011, Liaoning, China.
Omran Al-TosDepartment of Stomatology, The First Affiliated Hospital of Dalian Medical University, No.222, Zhongshan Road, Dalian, 116011, Liaoning, PR China.
Gang ChenDepartment of Stomatology, The First Affiliated Hospital of Dalian Medical University, No.222, Zhongshan Road, Dalian, 116011, Liaoning, PR China. 311121x@163.com.

Funding

Foundation of Liaoning Province Education Administration LJKZ0855The National Natural Science Foundation of China 81600818
6 · The paper itself

Abstract

backgroundArtificial intelligence AI platforms, such as Gemini, ChatGPT, DeepSeek, and Perplexity, are increasingly utilized to support clinical decision-making, yet their accuracy in specific medical domains remains variable. This study assessed the performance of these AI chatbots in responding to clinical questions commonly posed by surgeons in the context of salivary gland cancer, a field closely related to oral and maxillofacial oncology.

methodsThirty clinical questions related to salivary gland malignancies were created according to the ASCO 2021 guidelines. Two researchers posted on four AI chatbot platforms: ChatGPT-4o, DeepSeek, Gemini, and Peperlixity. These questions were queried three times daily over ten days, yielding a total of 2700 responses that were coded as correct or incorrect. The accuracy of each response was statistically analyzed, and overall accuracy rates for each AI platform were calculated.

resultsDeepSeek achieved the highest accuracy rate at 86.9%, followed by Gemini at 78.9%, ChatGPT-4o at 72.8%, and Perplexity at 71.6%.

conclusionDespite demonstrating substantial potential, current AI chatbots have not yet achieved sufficient accuracy for standalone clinical use in salivary gland cancer in clinical applications. Enhancements in AI capabilities and rigorous clinical validation are necessary to ensure patient safety and effectiveness in clinical practice.

Indexed as

Artificial IntelligenceClinical Decision-MakingSalivary Gland NeoplasmsGenerative Artificial IntelligenceHumansArtificial intelligenceChatGPTDeepSeekGeminiLLMsPerplexitySalivery gland cancer

Identifiers

PMID40849657
PMCPMC12374294

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.