ArticleBMC oral health2025
Evaluation of deepseek, gemini, ChatGPT-4o, and perplexity in responding to salivary gland cancer.
Article in BMC oral health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
11 citing papers in PubMed.
- Evaluation of large language model responses to patient questions on oral anticoagulant therapy: a comparative expert assessment.Exploratory research in clinical and social pharmacy · 2026Article
- Evaluating Large Reasoning Models Versus Human Multidisciplinary Teams in Lung Cancer Decision-Making: Real-World Study.Journal of medical Internet research · 2026Article
- Quality, completeness, readability, and response time of AI chatbots in the management of deep caries and pulp exposure: a bilingual comparative study.BMC oral health · 2026Article
- AI in respiratory care: findings from the GOLD report.Journal of translational medicine · 2026Article
- A comparative analysis of large language models for providing oral cavity cancer information.Scientific reports · 2026Article
- Ontology-driven generation of parameters for health technology assessment models: a prompt engineering study.International journal of technology assessment in health care · 2026Article
- Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit.BMJ open · 2026Article
- From algorithms to empathy: can large language models effectively answer patients' questions in restorative dentistry?BMC oral health · 2026Article
- Article
- Accuracy of large language models in head and neck cancers: a comparative analysis of ChatGPT and Gemini in TNM staging and clinical decision support.Frontiers in oncology · 2026Article
- Artificial Intelligence in Action: Racial and Gender Disparities in Academic Radiology.Cureus · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
Abstract
backgroundArtificial intelligence AI platforms, such as Gemini, ChatGPT, DeepSeek, and Perplexity, are increasingly utilized to support clinical decision-making, yet their accuracy in specific medical domains remains variable. This study assessed the performance of these AI chatbots in responding to clinical questions commonly posed by surgeons in the context of salivary gland cancer, a field closely related to oral and maxillofacial oncology.
methodsThirty clinical questions related to salivary gland malignancies were created according to the ASCO 2021 guidelines. Two researchers posted on four AI chatbot platforms: ChatGPT-4o, DeepSeek, Gemini, and Peperlixity. These questions were queried three times daily over ten days, yielding a total of 2700 responses that were coded as correct or incorrect. The accuracy of each response was statistically analyzed, and overall accuracy rates for each AI platform were calculated.
resultsDeepSeek achieved the highest accuracy rate at 86.9%, followed by Gemini at 78.9%, ChatGPT-4o at 72.8%, and Perplexity at 71.6%.
conclusionDespite demonstrating substantial potential, current AI chatbots have not yet achieved sufficient accuracy for standalone clinical use in salivary gland cancer in clinical applications. Enhancements in AI capabilities and rigorous clinical validation are necessary to ensure patient safety and effectiveness in clinical practice.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.