Evidence map›Paper›PMID 40973108›Full record

SynthesisJMIR medical education2025

Evaluating the Potential and Accuracy of ChatGPT-3.5 and 4.0 in Medical Licensing and In-Training Examinations: Systematic Review and Meta-Analysis.

Anila Jaleel, Umair Aziz, Ghulam Farid, Muhammad Zahid Bashir, Tehmasp Rehman Mirza, Syed Mohammad Khizar Abbas, Shiraz Aslam, Rana Muhammad Hassaan Sikander

Abstract readSystematic ReviewMeta-Analysis
In one paragraph

Synthesis in JMIR medical education, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 13 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
13citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

13 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Comment on "Experts V/S AI´s 2.0: comparative evaluation of ai models and expert consensus in obstructive sleep apnea assessment".European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Article
  5. Article
  6. Article
  7. Article
  8. ChatGPT in precision medicine.APL bioengineering · 2026
    Review
  9. Review
  10. Article
  11. Systematic Mining of Bioactive Compounds for Wound Healing FromJMIR bioinformatics and biotechnology · 2026
    Article
  12. Article
  13. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Anila Jaleel *Shalamar Medical and Dental College, Lahore, Pakistan.ORCID 0000-0002-5530-3696
Umair Aziz *Shalamar Medical and Dental College, Lahore, Pakistan.ORCID 0009-0008-7047-8858
Ghulam Farid *Shalamar Medical and Dental College, Lahore, Pakistan.ORCID 0000-0002-3299-5220
Muhammad Zahid BashirShalamar Medical and Dental College, Lahore, Pakistan.ORCID 0000-0003-4811-5782
Tehmasp Rehman MirzaShalamar Medical and Dental College, Lahore, Pakistan.ORCID 0009-0002-8572-4579
Syed Mohammad Khizar AbbasShalamar Medical and Dental College, Lahore, Pakistan.ORCID 0009-0004-9932-6779
Shiraz Aslam *Shalamar Medical and Dental College, Lahore, Pakistan.ORCID 0009-0006-4134-1890
Rana Muhammad Hassaan Sikander *Shalamar Medical and Dental College, Lahore, Pakistan.ORCID 0009-0003-3409-6223

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundArtificial intelligence (AI) has significantly impacted health care, medicine, and radiology, offering personalized treatment plans, simplified workflows, and informed clinical decisions. ChatGPT (OpenAI), a conversational AI model, has revolutionized health care and medical education by simulating clinical scenarios and improving communication skills. However, inconsistent performance across medical licensing examinations and variability between countries and specialties highlight the need for further research on contextual factors influencing AI accuracy and exploring its potential to enhance technical proficiency and soft skills, making AI a reliable tool in patient care and medical education.

objectiveThis systematic review aims to evaluate and compare the accuracy and potential of ChatGPT-3.5 and 4.0 in medical licensing and in-training residency examinations across various countries and specialties.

methodsA systematic review and meta-analysis were conducted, adhering to the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. Data were collected from multiple reputable databases (Scopus, PubMed, JMIR Publications, Elsevier, BMJ, and Wiley Online Library), focusing on studies published from January 2023 to July 2024. Analysis specifically targeted research assessing ChatGPT's efficacy in medical licensing exams, excluding studies not related to this focus or published in languages other than English. Ultimately, 53 studies were included, providing a robust dataset for comparing the accuracy rates of ChatGPT-3.5 and 4.0.

resultsChatGPT-4 outperformed ChatGPT-3.5 in medical licensing exams, achieving a pooled accuracy of 81.8%, compared to ChatGPT-3.5's 60.8%. In in-training residency exams, ChatGPT-4 achieved an accuracy rate of 72.2%, compared to 57.7% for ChatGPT-3.5. The forest plot presented a risk ratio of 1.36 (95% CI 1.30-1.43), demonstrating that ChatGPT-4 was 36% more likely to provide correct answers than ChatGPT-3.5 across both medical licensing and residency exams. These results indicate that ChatGPT-4 significantly outperforms ChatGPT-3.5, but the performance advantage varies depending on the exam type. This highlights the importance of targeted improvements and further research to optimize ChatGPT-4's performance in specific educational and clinical settings.

conclusionsChatGPT-4.0 and 3.5 show promising results in enhancing medical education and supporting clinical decision-making, but they cannot replace the comprehensive skill set required for effective medical practice. Future research should focus on improving AI's capabilities in interpreting complex clinical data and enhancing its reliability as an educational resource.

Indexed as

Artificial IntelligenceEducational MeasurementEducation, MedicalLicensure, MedicalClinical CompetenceGenerative Artificial IntelligenceHumansInternship and Residencyaccuracy of ChatGPTAI in health careartificial intelligenceChatGPTChatGPT-3.5 performanceChatGPT-4.0 performanceclinical decision-makingmedical educationmedical licensing examinations

Identifiers

PMID40973108
PMCPMC12495368

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.