ArticleJAMA ophthalmology2025
Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.
Article in JAMA ophthalmology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
15 citing papers in PubMed.
- Large Language Models Approximate Inter-Expert Agreement in Glaucoma Suspect and Glaucoma Classification from Multimodal Data.Ophthalmology science · 2026Article
- Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie · 2026Article
- Rethinking scale in ophthalmic artificial intelligence: from bigger models to smarter clinical reasoning.NPJ digital medicine · 2026Review
- Comparative Performance of Gemini 3 Pro and GPT-5 Family Models on Ophthalmology Board-Style Questions.Ophthalmology science · 2026Article
- Evaluating and enhancing the performance of large language models in thyroid eye disease through customization and Chain-of-Thought strategies.Scientific reports · 2026Article
- Performance of GPT-5 Frontier Models in Ophthalmology Question Answering.Ophthalmology science · 2026Article
- Benchmark evaluation of multi-modal large language models for ophthalmic diagnosis in real world.Frontiers in medicine · 2026Article
- ChatGPT-5 versus other mainstream large language models in core diabetic retinopathy patient queries.Frontiers in cell and developmental biology · 2026Article
- Evaluating large language models for diabetic retinopathy multiple-choice question generation in clinical ophthalmic education.Frontiers in medicine · 2026Article
- Reply.Ophthalmology science · 2026Article
- Benchmarking publicly accessible large language models for high-myopia multiple-choice question generation in digital ophthalmic education and public health training.Frontiers in public health · 2026Article
- Evaluating the Applicability of Advanced Large Language Models in Laboratory Medicine Test Questions: A Comparative Performance Study.Advances in medical education and practice · 2026Article
- Leveraging ChatGPT for Report Error Audit: An Accuracy-Driven and Cost-Efficient Solution for Ophthalmic Imaging Reports.Ophthalmology and therapy · 2025Article
- Comparing Artificial intelligence to physicians' competences in the domain of clinical reasoning: A systematic review and meta-analysis.Journal of medical education and curricular developmentReview
- Evaluating large language model clinical reasoning in glaucoma using retrieval-augmented generation.Advances in ophthalmology practice and researchArticle
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
19 authors.
Funding
Abstract
Importance: OpenAI's recent large language model (LLM) o1 has dedicated reasoning capabilities, but it remains untested in specialized medical fields like ophthalmology. Evaluating o1 in ophthalmology is crucial to determine whether its general reasoning can meet specialized needs or if domain-specific LLMs are warranted. Objective: To assess the performance and reasoning ability of OpenAI's o1 compared with other LLMs on ophthalmological questions. Design, Setting, and Participants: In September through October 2024, the LLMs o1, GPT-4o (OpenAI), GPT-4 (OpenAI), GPT-3.5 (OpenAI), Llama 3-8B (Meta), and Gemini 1.5 Pro (Google) were evaluated on 6990 standardized ophthalmology questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset. The study did not analyze human participants. Main Outcomes and Measures: Models were evaluated on performance (accuracy and macro F1 score) and reasoning abilities (text-generation metrics: Recall-Oriented Understudy for Gisting Evaluation [ROUGE-L], BERTScore, BARTScore, AlignScore, and Metric for Evaluation of Translation With Explicit Ordering [METEOR]). Mean scores are reported for o1, while mean differences (Δ) from o1's scores are reported for other models. Expert qualitative evaluation of o1 and GPT-4o responses assessed usefulness, organization, and comprehensibility using 5-point Likert scales. Results: The LLM o1 achieved the highest accuracy (mean, 0.877; 95% CI, 0.870 to 0.885) and macro F1 score (mean, 0.877; 95% CI, 0.869 to 0.884) (P < .001). In BERTScore, GPT-4o (Δ = 0.012; 95% CI, 0.012 to 0.013) and GPT-4 (Δ = 0.014; 95% CI, 0.014 to 0.015) outperformed o1 (P < .001). Similarly, in AlignScore, GPT-4o (Δ = 0.019; 95% CI, 0.016 to 0.021) and GPT-4 (Δ = 0.024; 95% CI, 0.021 to 0.026) again performed better (P < .001). In ROUGE-L, GPT-4o (Δ = 0.018; 95% CI, 0.017 to 0.019), GPT-4 (Δ = 0.026; 95% CI, 0.025 to 0.027), and GPT-3.5 (Δ = 0.008; 95% CI, 0.007 to 0.009) all outperformed o1 (P < .001). Conversely, o1 led in BARTScore (mean, -4.787; 95% CI, -4.813 to -4.762; P < .001) and METEOR (mean, 0.221; 95% CI, 0.218 to 0.223; P < .001 except GPT-4o). Also, o1 outperformed GPT-4o in usefulness (o1: mean, 4.81; 95% CI, 4.73 to 4.89; GPT-4o: mean, 4.53; 95% CI, 4.40 to 4.65; P < .001) and organization (o1: mean, 4.83; 95% CI, 4.75 to 4.90; GPT-4o: mean, 4.63; 95% CI, 4.51 to 4.74; P = .003). Conclusions and Relevance: This study found that o1 excelled in accuracy but showed inconsistencies in text-generation metrics, trailing GPT-4o and GPT-4; expert reviews found o1's responses to be more clinically useful and better organized than GPT-4o. While o1 demonstrated promise, its performance in addressing ophthalmology-specific challenges is not fully optimal, underscoring the potential need for domain-specialized LLMs and targeted evaluations.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.