ArticleOphthalmology science2026
BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for Knowledge and Reasoning.
Article in Ophthalmology science, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
1 citing paper in PubMed.
- Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
32 authors.
Funding
Abstract
Purpose: Current benchmarks evaluating large language models (LLMs) in ophthalmology are narrow and disproportionately prioritize accuracy. We introduce BEnchmarking LLMs for Ophthalmology (BELO), a standardized evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BEnchmarking LLMs for Ophthalmology assesses ophthalmology-related knowledge and reasoning quality. Subjects: This study did not involve human participation. Design: Cross-sectional study. Methods: Using keyword matching and a fine-tuned PubMed Bidirectional Encoder Representations from Transformers model, we curated ophthalmology-specific multiple-choice questions (MCQs) from diverse medical data sets (Basic and Clinical Science Course [BCSC], Multi-Subject Multi-Choice Dataset for Medical domain [MedMCQA], Medical Question Answering [MedQA], Biomedical Semantic Indexing and Question Answering [BioASQ], and PubMed Question Answering [PubMedQA]). The data set underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by 3 senior ophthalmologists. To illustrate BELO's utility, we evaluated 8 LLMs (OpenAI o1, o3-mini, GPT-5, GPT-4o, DeepSeek-R1, MedGemma-4B, Llama-3-8B, and Gemini 1.5 Pro). Main Outcome Measures: The 8 LLMs were evaluated in terms using accuracy, macro-F1, and 5 text-generation metrics (Recall-Oriented Understudy for Gisting Evaluation, BERTScore, BARTScore, Metric for Evaluation of Translation with Explicit Ordering, and AlignScore). In a further evaluation involving human experts, 2 ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. Results: BEnchmarking LLMs for Ophthalmology consists of 900 high-quality, expert-reviewed questions aggregated from 5 sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). To demonstrate BELO's utility, we conducted a series of benchmarking exercises. In the quantitative evaluation, GPT-5 achieved the highest accuracy (0.90, 95% confidence interval [CI]: 0.89-0.92) and macro-F1 score (0.91, 95% CI: 0.89-0.93). On the other hand, the models' performance on text-generation metrics varied and were generally suboptimal, with scores ranging from 20.4 to 72.0 (out of 100, excluding the BARTScore metric), indicating room for improvement in clinical reasoning. In expert evaluations, GPT-4o was rated highest for accuracy and readability, while Gemini 1.5 Pro scored highest for completeness. A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO data set will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models. Conclusions: BEnchmarking LLMs for Ophthalmology provides a robust clinically relevant benchmark for evaluating both the accuracy and reasoning capabilities of current and emerging LLMs in ophthalmology. Future BELO benchmarking efforts will expand to include vision-based question answering and clinical scenario management tasks. Financial Disclosures: The authors have no proprietary or commercial interest in any materials discussed in this article.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.