Evidence map›Paper›PMID 41696659›Full record

ArticleOphthalmology science2026

BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for Knowledge and Reasoning.

Sahana Srinivasan, Xuguang Ai, Thaddaeus Wai Soon Lo, Aidan Gilson, Minjie Zou, Ke Zou, Hyunjae Kim, Mingjia Yang, Krithi Pushpanathan, Samantha Min Er Yew and 22 more

Abstract read
In one paragraph

Article in Ophthalmology science, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

32 authors.

Sahana SrinivasanDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Xuguang AiDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven.
Thaddaeus Wai Soon LoSingapore Eye Research Institute, Singapore National Eye Centre, Singapore.
Aidan GilsonDepartment of Ophthalmology, Massachusetts Eye and Ear, Harvard Medical School, Boston, Massachusetts.
Minjie ZouDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Ke ZouCentre for Innovation and Precision Eye Health, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Hyunjae KimDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven.
Mingjia YangSchool of Chemistry, Chemical Engineering and Biotechnology, Nanyang Technological University, Singapore.
Krithi PushpanathanDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Samantha Min Er YewDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Wan Ting LokeDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Jocelyn Hui Lin GohDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Yibing ChenDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Yiming KongDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven.
Emily Yuelei FuDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven.
Michelle OngEngineering Product Development Pillar, Singapore University of Technology and Design, Singapore.
Kristen NwanyanwuDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven.
Amisha DaveDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven.
Kelvin Zhenghao LiDepartment of Ophthalmology, Tan Tock Seng Hospital, Singapore.
Chen-Hsin SunDepartment of Ophthalmology, National University Hospital, Singapore.
Mark ChiaNIHR Biomedical Research Centre at Moorfields Eye Hospital NHS Foundation Trust, London, UK.
Gabriel Dawei YangSingapore Eye Research Institute, Singapore National Eye Centre, Singapore.
Wendy Meihua WongDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
David Ziyou ChenDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Dianbo LiuDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.
Maxwell SingerDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven.
Fares AntakiCole Eye Institute, Cleveland Clinic, Cleveland, Ohio.
Lucian V Del PrioreDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven.
Jost B JonasSingapore Eye Research Institute, Singapore National Eye Centre, Singapore.
Ron AdelmanDepartment of Ophthalmology and Visual Science, Yale School of Medicine, Yale University, New Haven.
Qingyu ChenDepartment of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven.
Yih-Chung ThamDepartment of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore.

Funding

Addressing Factual Inaccuracy and Unfaithful Reasoning of Large Language Models in Biomedicine and HealthcareR01LM014604 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$1.1M
Natural language processing and medical imaging analysis for multi-modality computer assisted diagnosis of ophthalmic diseasesR00LM014024 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$747k
NLM NIH HHS R00 LM014024NLM NIH HHS R01 LM014604
6 · The paper itself

Abstract

Purpose: Current benchmarks evaluating large language models (LLMs) in ophthalmology are narrow and disproportionately prioritize accuracy. We introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BEnchmarking LLMs for Ophthalmology assesses ophthalmology-related knowledge and reasoning quality. Subjects: This study did not involve human participation. Design: Cross-sectional study. Methods: Using keyword matching and a fine-tuned PubMed Bidirectional Encoder Representations from Transformers model, we curated ophthalmology-specific multiple-choice questions (MCQs) from diverse medical data sets (Basic and Clinical Science Course [BCSC], Multi-Subject Multi-Choice Dataset for Medical domain [MedMCQA], Medical Question Answering [MedQA], Biomedical Semantic Indexing and Question Answering [BioASQ], and PubMed Question Answering [PubMedQA]). The data set underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by 3 senior ophthalmologists. To illustrate BELO's utility, we evaluated 8 LLMs (OpenAI o1, o3-mini, GPT-5, GPT-4o, DeepSeek-R1, MedGemma-4B, Llama-3-8B, and Gemini 1.5 Pro). Main Outcome Measures: The 8 LLMs were evaluated in terms using accuracy, macro-F1, and 5 text-generation metrics (Recall-Oriented Understudy for Gisting Evaluation, BERTScore, BARTScore, Metric for Evaluation of Translation with Explicit Ordering, and AlignScore). In a further evaluation involving human experts, 2 ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. Results: BEnchmarking LLMs for Ophthalmology consists of 900 high-quality, expert-reviewed questions aggregated from 5 sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). To demonstrate BELO's utility, we conducted a series of benchmarking exercises. In the quantitative evaluation, GPT-5 achieved the highest accuracy (0.90, 95% confidence interval [CI]: 0.89-0.92) and macro-F1 score (0.91, 95% CI: 0.89-0.93). On the other hand, the models' performance on text-generation metrics varied and were generally suboptimal, with scores ranging from 20.4 to 72.0 (out of 100, excluding the BARTScore metric), indicating room for improvement in clinical reasoning. In expert evaluations, GPT-4o was rated highest for accuracy and readability, while Gemini 1.5 Pro scored highest for completeness. A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO data set will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models. Conclusions: BEnchmarking LLMs for Ophthalmology provides a robust clinically relevant benchmark for evaluating both the accuracy and reasoning capabilities of current and emerging LLMs in ophthalmology. Future BELO benchmarking efforts will expand to include vision-based question answering and clinical scenario management tasks. Financial Disclosures: The authors have no proprietary or commercial interest in any materials discussed in this article.

Indexed as

Benchmark data setExpert curatedLarge language modelsOphthalmological reasoningQuestion–answer

Identifiers

PMID41696659
PMCPMC12906013

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.