Evidence map›Paper›PMID 41727573›Full record

ArticleResearch square2026

How Far Have Large Language Models Advanced in Ophthalmology? A Systematic Review of Their Development, Evaluation, and Readiness for Clinical Use.

Hyunjae Kim, Yu Yin, Zhiyuan Cao, Chen Liu, Anran Li, Zhen Chen, Xuguang Ai, Younjoon Chung, Fan Ma, Xueping Peng and 11 more

Abstract readPreprint
In one paragraph

Article in Research square, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

21 authors.

Hyunjae KimYale University.
Yu YinUniversity of Queensland.
Zhiyuan CaoYale University.
Chen LiuYale University.
Anran LiYale University.
Zhen ChenYale University.
Xuguang AiYale University.
Younjoon ChungYale University.
Fan MaYale University.
Xueping PengYale University.
Lingfei QianYale University.
Zhenyue QinYale University.
Kalpana RajaYale University.
Yang RenYale University.
Weipeng ZhouYale University.
Yih-Chung ThamNational University of Singapore.
Emily Y ChewNational Institutes of Health.
Zhiyong LuNational Institutes of Health.
Sophia Y WangStanford University.
Hua XuYale University.
Qingyu ChenYale University.

Funding

Addressing Factual Inaccuracy and Unfaithful Reasoning of Large Language Models in Biomedicine and HealthcareR01LM014604 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$1.1M
Natural language processing and medical imaging analysis for multi-modality computer assisted diagnosis of ophthalmic diseasesR00LM014024 · NLM · YALE UNIVERSITY · PI Qingyu Chen · 2024 to 2026
$747k
NLM NIH HHS R00 LM014024NLM NIH HHS R01 LM014604
6 · The paper itself

Abstract

Large language models (LLMs) are rapidly transforming ophthalmology, with expanding applications in patient care, clinical documentation, and medical education. Recent studies span a wide range of use cases-from early text-only applications to emerging multimodal systems that integrate ophthalmic images to support diagnosis and generate assessment and treatment plans. Amid this rapid progress, it is critical for both researchers and clinicians to stay informed in order to guide responsible development and adoption. However, prior reviews have largely focused on narrow domains such as an inventory of potential use cases or performance on board-style examinations, leaving the broader landscape insufficiently characterized. Key questions remain unanswered: How are LLMs in ophthalmology being developed? What applications and evaluation strategies are being pursued? And which areas are closest to real-world clinical adoption? To date, these aspects have not been comprehensively examined. In this study, we conducted a systematic review on LLMs in ophthalmology by manually screening 1, 029 studies from PubMed/PMC, Scopus, and Embase published between January 1, 2022, and April 1, 2025, identifying 91 relevant articles. To provide a standardized assessment, we introduced a structured framework that categorizes ophthalmic use cases and stratifies evaluation rigor across five levels of maturity. Each study was manually annotated using 27 structured variables spanning multiple dimensions: scope and purpose (e.g., study aim, ophthalmic subspecialty, input modality); model architecture and training (e.g., backbone LLMs, domain-specific adaptations); evaluation and validation (e.g., target applications, evaluation metrics, level of clinical validation); and resource availability (e.g., model access, licensing, dataset availability). We additionally performed a small-scale, illustrative evaluation of representative emerging models, such as GPT-5.2, gpt-oss-120B, and Gemini 3, to contextualize previously reported results on commonly used ophthalmology tasks. The results show that most studies focused on general-purpose proprietary models, such as GPT-4 and Gemini, while fewer than 10% introduced domain-specific adaptations for ophthalmology, including only 4% that developed ophthalmology-specific architectures for text-based applications. Multimodal LLMs remain relatively underexplored, with only 23% of studies incorporating imaging data. Evaluation practices reveal a significant translational gap: While 57.1% of studies relied on standard benchmarking and expert review, only 9.9% conducted retrospective validation using real-world clinical data, and just two studies progressed to prospective pilot evaluation. Moreover, although model performance on benchmarks on board-style exams and clinical vignettes has improved with newer model generations, reproducibility and transparency remain limited: only 5.5% of studies released evaluation code, and 33% used publicly available datasets. Finally, we provide a living repository to track the rapid progress of LLMs in ophthalmology for the broader research and clinical community.

Identifiers

PMID41727573
PMCPMC12919167

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.