Evidence map›Paper›PMID 40742581›Full record

ArticleJAMA ophthalmology2025

Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.

Sahana Srinivasan, Xuguang Ai, Minjie Zou, Ke Zou, Hyunjae Kim, Thaddaeus Wai Soon Lo, Krithi Pushpanathan, Gabriel Dawei Yang, Jocelyn Hui Lin Goh, Yiming Kong and 9 more

Abstract read
In one paragraph

Article in JAMA ophthalmology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.

0numbers the graph read from it
0cells of the map it votes in
15citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

15 citing papers in PubMed.

  1. Article
  2. Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie · 2026
    Article
  3. Review
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Reply.Ophthalmology science · 2026
    Article
  11. Article
  12. Article
  13. Article
  14. Review
  15. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

19 authors.

Sahana SrinivasanCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Xuguang AiDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Minjie ZouCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Ke ZouCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Hyunjae KimDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Thaddaeus Wai Soon LoSingapore Eye Research Institute, Singapore National Eye Centre, Singapore.
Krithi PushpanathanCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Gabriel Dawei YangSingapore Eye Research Institute, Singapore National Eye Centre, Singapore.
Jocelyn Hui Lin GohCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Yiming KongDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Anran LiDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Maxwell B SingerDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Kai JinEye Center, The Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, Zhejiang, China.
Fares AntakiCole Eye Institute, Cleveland Clinic, Cleveland, Ohio.
David Ziyou ChenCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Dianbo LiuCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Ron A AdelmanDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Qingyu ChenDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, Connecticut.
Yih Chung ThamCentre for Innovation and Precision Eye Health, Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.

Funding

Addressing Factual Inaccuracy and Unfaithful Reasoning of Large Language Models in Biomedicine and HealthcareR01LM014604 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$1.1M
Natural language processing and medical imaging analysis for multi-modality computer assisted diagnosis of ophthalmic diseasesR00LM014024 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$747k
Natural language processing and medical imaging analysis for multi-modality computer assisted diagnosis of ophthalmic diseasesK99LM014024 · NLM · YALE UNIVERSITY · PI CHEN, QINGYU · 2023 to 2023
$92k
NLM NIH HHS K99 LM014024NLM NIH HHS R00 LM014024NLM NIH HHS R01 LM014604
6 · The paper itself

Abstract

Importance: OpenAI's recent large language model (LLM) o1 has dedicated reasoning capabilities, but it remains untested in specialized medical fields like ophthalmology. Evaluating o1 in ophthalmology is crucial to determine whether its general reasoning can meet specialized needs or if domain-specific LLMs are warranted. Objective: To assess the performance and reasoning ability of OpenAI's o1 compared with other LLMs on ophthalmological questions. Design, Setting, and Participants: In September through October 2024, the LLMs o1, GPT-4o (OpenAI), GPT-4 (OpenAI), GPT-3.5 (OpenAI), Llama 3-8B (Meta), and Gemini 1.5 Pro (Google) were evaluated on 6990 standardized ophthalmology questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset. The study did not analyze human participants. Main Outcomes and Measures: Models were evaluated on performance (accuracy and macro F1 score) and reasoning abilities (text-generation metrics: Recall-Oriented Understudy for Gisting Evaluation [ROUGE-L], BERTScore, BARTScore, AlignScore, and Metric for Evaluation of Translation With Explicit Ordering [METEOR]). Mean scores are reported for o1, while mean differences (Δ) from o1's scores are reported for other models. Expert qualitative evaluation of o1 and GPT-4o responses assessed usefulness, organization, and comprehensibility using 5-point Likert scales. Results: The LLM o1 achieved the highest accuracy (mean, 0.877; 95% CI, 0.870 to 0.885) and macro F1 score (mean, 0.877; 95% CI, 0.869 to 0.884) (P < .001). In BERTScore, GPT-4o (Δ = 0.012; 95% CI, 0.012 to 0.013) and GPT-4 (Δ = 0.014; 95% CI, 0.014 to 0.015) outperformed o1 (P < .001). Similarly, in AlignScore, GPT-4o (Δ = 0.019; 95% CI, 0.016 to 0.021) and GPT-4 (Δ = 0.024; 95% CI, 0.021 to 0.026) again performed better (P < .001). In ROUGE-L, GPT-4o (Δ = 0.018; 95% CI, 0.017 to 0.019), GPT-4 (Δ = 0.026; 95% CI, 0.025 to 0.027), and GPT-3.5 (Δ = 0.008; 95% CI, 0.007 to 0.009) all outperformed o1 (P < .001). Conversely, o1 led in BARTScore (mean, -4.787; 95% CI, -4.813 to -4.762; P < .001) and METEOR (mean, 0.221; 95% CI, 0.218 to 0.223; P < .001 except GPT-4o). Also, o1 outperformed GPT-4o in usefulness (o1: mean, 4.81; 95% CI, 4.73 to 4.89; GPT-4o: mean, 4.53; 95% CI, 4.40 to 4.65; P < .001) and organization (o1: mean, 4.83; 95% CI, 4.75 to 4.90; GPT-4o: mean, 4.63; 95% CI, 4.51 to 4.74; P = .003). Conclusions and Relevance: This study found that o1 excelled in accuracy but showed inconsistencies in text-generation metrics, trailing GPT-4o and GPT-4; expert reviews found o1's responses to be more clinically useful and better organized than GPT-4o. While o1 demonstrated promise, its performance in addressing ophthalmology-specific challenges is not fully optimal, underscoring the potential need for domain-specialized LLMs and targeted evaluations.

Indexed as

LanguageOphthalmologyHumansLarge Language Models

Identifiers

PMID40742581
PMCPMC12314776

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.