Evidence map›Paper›PMID 41158456›Full record

ArticleFrontiers in medicine2025

Evaluating AI performance in infectious disease education: a comparative analysis of ChatGPT, Google Bard, Perplexity AI, Microsoft Copilot, and Meta AI.

Abdulaziz Ibrahim Alzarea, Azfar Athar Ishaqui, Muhammad Bilal Maqsood, Abdullah Salah Alanazi, Aseel Awad Alsaidan, Tauqeer Hussain Mallhi, Narendar Kumar, Muhammad Imran, Sultan M Alshahrani, Hassan H Alhassan and 2 more

Abstract read
In one paragraph

Article in Frontiers in medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Abdulaziz Ibrahim Alzarea *Department of Clinical Pharmacy, College of Pharmacy, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.
Azfar Athar Ishaqui *Department of Clinical Pharmacy, College of Pharmacy, King Khalid University, Abha, Saudi Arabia.
Muhammad Bilal MaqsoodEastern Health Cluster, Ministry of Health, Dammam, Saudi Arabia.
Abdullah Salah AlanaziDepartment of Clinical Pharmacy, College of Pharmacy, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.
Aseel Awad AlsaidanDepartment of Family and Community Medicine, College of Medicine, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.
Tauqeer Hussain MallhiMedicines R Us Chemist, Gregory Hills, NSW, Australia.
Narendar KumarDepartment of Pharmacy Practice, Faculty of Pharmacy, Sindh University, Jamshoro, Pakistan.
Muhammad ImranDepartment of Pharmacy, Iqra University, Karachi, Pakistan.
Sultan M AlshahraniDepartment of Clinical Pharmacy, College of Pharmacy, King Khalid University, Abha, Saudi Arabia.
Hassan H AlhassanDepartment of Clinical Laboratory Sciences, College of Applied Medical Sciences, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.
Sami I AlzareaDepartment of Pharmacology, College of Pharmacy, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.
Omar Awad AlsaidanDepartment of Pharmaceutics, College of Pharmacy, Jouf University, Sakaka, Al-Jouf, Saudi Arabia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: This study systematically evaluates and compares the performance of ChatGPT 3. 5, Google Bard (Gemini), Perplexity AI, Microsoft Copilot, and Meta AI in responding to infectious disease-related multiple-choice questions (MCQs). Methods: A systematic comparative study was conducted using 20 infectious disease case studies sourced from Infectious Diseases: A Case Study Approach by Jonathan C. Cho. Each case study included 7-10 MCQs, resulting in a total of 160 questions. AI platforms were provided with standardized prompts containing the case study text and MCQs without additional context. Their responses were evaluated against a reference answer key from the textbook. Accuracy was measured by the percentage of correct responses, and consistency was assessed by submitting identical prompts 24 h apart. Results: ChatGPT 3.5 achieved the highest numerical accuracy (65.6%), followed by Perplexity AI (63.2%), Microsoft Copilot (60.9%), Meta AI (60.8%), and Google Bard (58.8%). AI models performed best in symptom identification (76.5%) and worst in therapy-related questions (57.1%). ChatGPT 3.5 demonstrated strong diagnostic accuracy (79.1%) but had a significant drop in antimicrobial treatment recommendations (56.6%). Google Bard performed inconsistently in microorganism identification (61.9%) and preventive therapy (62.5%). Microsoft Copilot exhibited the most stable responses across repeated testing, while ChatGPT 3.5 showed a 7.5% accuracy decline. Perplexity AI and Meta AI struggled with individualized treatment recommendations, showing variability in drug selection and dosing adjustments. AI-generated responses were found to change over time, with some models giving different antimicrobial recommendations for the same case scenario upon repeated testing. Conclusion: AI platforms offer potential in infectious disease education but demonstrate limitations in pharmacotherapy decision-making, particularly in antimicrobial selection and dosing accuracy. ChatGPT 3.5 performed best but lacked response stability, while Microsoft Copilot showed greater consistency but lacked nuanced therapeutic reasoning. Further research is needed to improve AI-driven decision support systems for medical education and clinical applications through clinical trials, evaluation of real-world patient data, and assessment of long-term stability.

Indexed as

artificial intelligenceChatGPTGoogle Bardinfectious diseaseMeta AIMicrosoft CopilotPerplexity AI

Identifiers

PMID41158456
PMCPMC12557576

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.