Evidence map›Paper›PMID 41144954›Full record

SynthesisJournal of medical Internet research2025

Evaluating Large Language Models in Ophthalmology: Systematic Review.

Zili Zhang, Haiyang Zhang, Zhe Pan, Zhangqian Bi, Yao Wan, Xuefei Song, Xianqun Fan

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 16 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
16citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

16 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Review
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Review
  14. Article
  15. Article
  16. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Zili Zhang *State Key Laboratory of Eye Health, Department of Ophthalmology, Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.ORCID https://orcid.org/0009-0001-9259-666X
Haiyang Zhang *State Key Laboratory of Eye Health, Department of Ophthalmology, Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.ORCID https://orcid.org/0009-0001-2365-7609
Zhe PanSchool of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.ORCID https://orcid.org/0009-0005-4402-5511
Zhangqian BiSchool of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.ORCID https://orcid.org/0000-0003-2257-9052
Yao WanSchool of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.ORCID https://orcid.org/0000-0001-6937-4180
Xuefei SongState Key Laboratory of Eye Health, Department of Ophthalmology, Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.ORCID https://orcid.org/0000-0002-5666-6894
Xianqun FanState Key Laboratory of Eye Health, Department of Ophthalmology, Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.ORCID https://orcid.org/0000-0002-9394-3969

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundLarge language models (LLMs) have the potential to revolutionize ophthalmic care, but their evaluation practice remains fragmented. A systematic assessment is crucial to identify gaps and guide future evaluation practices and clinical integration.

objectiveThis study aims to map the current landscape of LLM evaluations in ophthalmology and explore whether performance synthesis is feasible for a common task.

methodsA comprehensive search of PubMed, Web of Science, Embase, and IEEE Xplore was conducted up to November 17, 2024 (no language limits). Eligible publications quantitatively assessed an existing or modified LLM on ophthalmology-related tasks. Studies without full-text availability or those focusing solely on vision-only models were excluded. Two reviewers screened studies and extracted data across 6 dimensions (evaluated LLM, data modality, ophthalmic subspecialty, medical task, evaluation dimension, and clinical alignment), and disagreements were resolved by a third reviewer. Descriptive statistics were analyzed and visualized using Python (with NumPy, Pandas, SciPy, and Matplotlib libraries). The Fisher exact test compared open- versus closed-source models. An exploratory random-effects meta-analysis (logit transformation; DerSimonian-Laird τ

resultsOf the 817 identified records, 187 studies met the inclusion criteria. Closed-source LLMs dominated: 170 for ChatGPT, 58 for Gemini, and 32 for Copilot. Open-source LLMs appeared in only 25 (13.4%) of studies overall, but they appeared in 17 (77.3%) of evaluation-after-development studies, versus 8 (4.8%) pure-evaluation studies (P<1×10

conclusionsEvidence on LLM evaluations in ophthalmology is extensive but heterogeneous. Most studies have tested a few closed-source LLMs on text-based questions, leaving open-source systems, multimodal tasks, non-English contexts, and real-world deployment underexamined. High methodological variability precludes meaningful performance aggregation, as illustrated by the heterogeneous meta-analysis. Standardized, multimodal benchmarks and phased clinical validation pipelines are urgently needed before LLMs can be safely integrated into eye care workflows.

Indexed as

LanguageOphthalmologyHumansLarge Language Modelsartificial intelligenceclinical evaluationlarge language modelmeta-analysisophthalmologysystematic review

Identifiers

PMID41144954
PMCPMC12603593

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.