SynthesisJMIR medical education2025
Evaluating the Potential and Accuracy of ChatGPT-3.5 and 4.0 in Medical Licensing and In-Training Examinations: Systematic Review and Meta-Analysis.
Synthesis in JMIR medical education, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 13 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
13 citing papers in PubMed, 1 synthesis or guideline pooled it.
- The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis.Anatomical sciences education · 2026Pooled it
- Comparative evaluation of handwriting recognition by large language models (LLMs) in interpreting handwritten medication lists.Exploratory research in clinical and social pharmacy · 2026Article
- Enhancing Psychiatry Training Using an Agentic AI Simulated Consultation Tool: Prospective Cohort Study.JMIR medical education · 2026Article
- Comment on "Experts V/S AI´s 2.0: comparative evaluation of ai models and expert consensus in obstructive sleep apnea assessment".European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026Article
- Performance of large language models on the Turkish Pharmacy Specialty Examination: a comparative analysis of accuracy, confidence, and readability.Scientific reports · 2026Article
- Treatment Regret in Patients Undergoing Minimally Invasive Treatments for Benign Prostatic Hyperplasia.Journal of clinical medicine · 2026Article
- Artificial Intelligence in Psychiatry Training: Comparative Insights from Nine Large Language Models Across Cultural and Exam Contexts.The Psychiatric quarterly · 2026Article
- ChatGPT in precision medicine.APL bioengineering · 2026Review
- Evaluation of ChatGPT's Accuracy, Repeatability, and Reasoning Ability in Prosthodontics Education: A Cross-Sectional Comparative Study with Prosthodontists.Journal of clinical and experimental dentistry · 2026Review
- AI-Driven Objective Structured Clinical Examination Generation in Digital Health Education: Comparative Analysis of Three GPT-4o Configurations.JMIR medical education · 2026Article
- Systematic Mining of Bioactive Compounds for Wound Healing FromJMIR bioinformatics and biotechnology · 2026Article
- AI-assisted vs. textbook-based vs. blended learning for acute abdomen diagnosis: a retrospective cohort study of emergency interns.Frontiers in public health · 2026Article
- Feasibility of Real-time Artificial Intelligence-based Language Translation for Bilingual Informed Consent in Interventional Radiology:An Australian Proof-of-concept.Interventional radiology (Higashimatsuyama-shi (Japan) · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundArtificial intelligence (AI) has significantly impacted health care, medicine, and radiology, offering personalized treatment plans, simplified workflows, and informed clinical decisions. ChatGPT (OpenAI), a conversational AI model, has revolutionized health care and medical education by simulating clinical scenarios and improving communication skills. However, inconsistent performance across medical licensing examinations and variability between countries and specialties highlight the need for further research on contextual factors influencing AI accuracy and exploring its potential to enhance technical proficiency and soft skills, making AI a reliable tool in patient care and medical education.
objectiveThis systematic review aims to evaluate and compare the accuracy and potential of ChatGPT-3.5 and 4.0 in medical licensing and in-training residency examinations across various countries and specialties.
methodsA systematic review and meta-analysis were conducted, adhering to the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. Data were collected from multiple reputable databases (Scopus, PubMed, JMIR Publications, Elsevier, BMJ, and Wiley Online Library), focusing on studies published from January 2023 to July 2024. Analysis specifically targeted research assessing ChatGPT's efficacy in medical licensing exams, excluding studies not related to this focus or published in languages other than English. Ultimately, 53 studies were included, providing a robust dataset for comparing the accuracy rates of ChatGPT-3.5 and 4.0.
resultsChatGPT-4 outperformed ChatGPT-3.5 in medical licensing exams, achieving a pooled accuracy of 81.8%, compared to ChatGPT-3.5's 60.8%. In in-training residency exams, ChatGPT-4 achieved an accuracy rate of 72.2%, compared to 57.7% for ChatGPT-3.5. The forest plot presented a risk ratio of 1.36 (95% CI 1.30-1.43), demonstrating that ChatGPT-4 was 36% more likely to provide correct answers than ChatGPT-3.5 across both medical licensing and residency exams. These results indicate that ChatGPT-4 significantly outperforms ChatGPT-3.5, but the performance advantage varies depending on the exam type. This highlights the importance of targeted improvements and further research to optimize ChatGPT-4's performance in specific educational and clinical settings.
conclusionsChatGPT-4.0 and 3.5 show promising results in enhancing medical education and supporting clinical decision-making, but they cannot replace the comprehensive skill set required for effective medical practice. Future research should focus on improving AI's capabilities in interpreting complex clinical data and enhancing its reliability as an educational resource.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.