Evidence map›Paper›PMID 41133609›Full record

ArticleVision (Basel, Switzerland)2025

Comparative Assessment of Large Language Models in Optics and Refractive Surgery: Performance on Multiple-Choice Questions.

Leah Attal, Elad Shvartz, Alon Gorenshtein, Shirley Pincovich, Daniel Bahir

Abstract read
In one paragraph

Article in Vision (Basel, Switzerland), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Leah AttalAzrieli Faculty of Medicine, Bar Ilan University, Ramat-Gan 5290002, Israel.ORCID 0009-0000-0166-0631
Elad ShvartzAzrieli Faculty of Medicine, Bar Ilan University, Ramat-Gan 5290002, Israel.ORCID 0009-0007-9171-7996
Alon GorenshteinAzrieli Faculty of Medicine, Bar Ilan University, Ramat-Gan 5290002, Israel.ORCID 0009-0000-7542-8608
Shirley PincovichAzrieli Faculty of Medicine, Bar Ilan University, Ramat-Gan 5290002, Israel.
Daniel BahirAzrieli Faculty of Medicine, Bar Ilan University, Ramat-Gan 5290002, Israel.ORCID 0000-0002-1232-6257

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This study aimed to evaluate the performance of seven advanced AI Large Language Models (LLMs)-ChatGPT 4o, ChatGPT O3 Mini, ChatGPT O1, DeepSeek V3, DeepSeek R1, Gemini 2.0 Flash, and Grok-3-in answering multiple-choice questions (MCQs) in optics and refractive surgery, to assess their role in medical education for residents. The AI models were tested using 134 publicly available MCQs from national ophthalmology certification exams, categorized by the need to perform calculations, the relevant subspecialty, and the use of images. Accuracy was analyzed and compared statistically. ChatGPT O1 achieved the highest overall accuracy (83.5%), excelling in complex optical calculations (84.1%) and optics questions (82.4%). DeepSeek V3 displayed superior accuracy in refractive surgery-related questions (89.7%), followed by ChatGPT O3 Mini (88.4%). ChatGPT O3 Mini significantly outperformed others in image analysis, with 88.2% accuracy. Moreover, ChatGPT O1 demonstrated comparable accuracy rates for both calculated and non-calculated questions (84.1% vs. 83.3%). This is in stark contrast to other models, which exhibited significant discrepancies in accuracy for calculated and non-calculated questions. The findings highlight the ability of LLMs to achieve high accuracy in ophthalmology MCQs, particularly in complex optical calculations and visual items. These results suggest potential applications in exam preparation and medical training contexts, while underscoring the need for future studies designed to directly evaluate their role and impact in medical education. The findings highlight the significant potential of AI models in ophthalmology education, particularly in performing complex optical calculations and visual stem questions. Future studies should utilize larger, multilingual datasets to confirm and extend these preliminary findings.

Indexed as

artificial intelligenceChatGPTlarge language modelsmedical educationmultiple-choice questionsoptics

Identifiers

PMID41133609
PMCPMC12550897

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.