SynthesisJournal of medical systems2025
Evaluating the Performance of ChatGPT on Board-Style Examination Questions in Ophthalmology: A Meta-Analysis.
Synthesis in Journal of medical systems, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
15 citing papers in PubMed, 1 synthesis or guideline pooled it.
- The performance of ChatGPT and other large language models on multiple-choice questions in biomedical disciplines: A meta-analysis.Anatomical sciences education · 2026Pooled it
- Large Language Models for Ophthalmology Training in China: A Prospective Evaluation.Ophthalmology science · 2026Article
- Comparing the performance of four mainstream large language models on medical literature review generation: a human expert evaluation in SMILE surgery.Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie · 2026Article
- How Far Have Large Language Models Advanced in Ophthalmology? A Systematic Review of Their Development, Evaluation, and Readiness for Clinical Use.Research square · 2026Article
- Application effect and teaching evaluation of case-based learning combined with ChatGPT in ophthalmology clinical teaching.Frontiers in medicine · 2026Article
- Patient education for neuromyelitis optica spectrum disorder using large language models: combining expert assessment and real-world patient interaction.Frontiers in physiology · 2026Article
- Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering.Frontiers in cell and developmental biology · 2026Article
- ChatGPT-5 versus other mainstream large language models in core diabetic retinopathy patient queries.Frontiers in cell and developmental biology · 2026Article
- Benchmarking large language models for congenital cataract parent counseling: safety, readability, and knowledge translation of developmental and genetic information.Frontiers in cell and developmental biology · 2026Article
- Questionnaire on efficacy of the competency-oriented integrated residency and fellowship training for ophthalmologists in Shanghai.Frontiers in medicine · 2026Article
- Benchmark evaluation of multi-modal large language models for ophthalmic diagnosis in real world.Frontiers in medicine · 2026Article
- Review
- Comparative performance of large language models for patient-initiated ophthalmology consultations.Frontiers in public health · 2025Article
- Advances in the application of artificial intelligence in ophthalmic education and clinical training.Frontiers in medicine · 2025Review
- Evaluating large language model clinical reasoning in glaucoma using retrieval-augmented generation.Advances in ophthalmology practice and researchArticle
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
To review empirical research on ChatGPT's accuracy in answering ophthalmology board-style examination questions up to March 2025 and to analyze the effects of GPT versions, question types, language differences, and ophthalmology topics on accuracy. A search was conducted in PubMed, Web of Science, Embase, Scopus, and the Cochrane Library in March 2025. Two authors extracted data and independently assessed study quality. Accuracy rates were calculated with Stata 17.0. GPT-4 had an integrated accuracy of 73%, higher than GPT-3.5's 54%. It scored 77% in text and 55% in image tasks. GPT-4's accuracy was 73% in English-speaking countries and 71% in non-English ones. In ophthalmology, General Medicine achieved the highest accuracy (80%), while Clinical Optics had the lowest performance (55%). GPT-4 outperforms GPT-3.5, but its image processing capability needs further validation. Performance varies by language and topic, suggesting the need for more research on cross-linguistic efficacy and error analysis.
Indexed as
Identifiers
40615678What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.