ArticleOphthalmology science
Glaucoma Detection and Feature Identification via GPT-4V Fundus Image Analysis.
Article in Ophthalmology science. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
15 citing papers in PubMed.
- Image-Quality-Aware Multimodal Artificial Intelligence for Automated Structured OCT Report Generation in Glaucoma Evaluation.Ophthalmology science · 2026Article
- Hypertuned boosting approach with Local Binary Pattern and Pivot Distribution Count method feature extractor for glaucoma identification.Scientific reports · 2026Article
- Comprehensive Evaluation of ChatGPT's Diagnostic Accuracy on Image-based Ophthalmic Case Interpretations.Ophthalmology science · 2026Article
- A comparison of GPT-4V's capability in optical coherence tomography images of age-related macular degeneration with expert assessments.BMC ophthalmology · 2026Article
- Multimodal Diagnostic Accuracy of GPT-4.o, Claude 3.7, and Gemini 2.5 on Real-World Retina Cases.Journal of vitreoretinal diseases · 2026Article
- ChatGPT-Assisted Glaucoma Diagnosis: A Health-Equitable Multi-Ancestry Analysis Using Visual Field and Optical Coherence Tomography Data.American journal of ophthalmology · 2026Article
- How Far Have Large Language Models Advanced in Ophthalmology? A Systematic Review of Their Development, Evaluation, and Readiness for Clinical Use.Research square · 2026Article
- Applications of Large Language Models in Glaucoma: A Scoping Review.Vision (Basel, Switzerland) · 2026Review
- Can Multimodal Large Language Models Diagnose Diabetic Retinopathy from Fundus Photos? A Quantitative Evaluation.Ophthalmology science · 2026Article
- Large language models for ophthalmic examination understanding: from information extraction to clinical decision support.Frontiers in medicine · 2026Review
- Evaluation of DeepSeek-R1 for Ophthalmic Diagnosis and Reasoning: A Comparison with OpenAI o1 and o3.Journal of medical systems · 2025Article
- Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model.ArXiv · 2025Article
- Evaluating the clinical utility of multimodal large language models for detecting age-related macular degeneration from retinal imaging.Scientific reports · 2025Article
- Augmented Decisions: AI-Enhanced Accuracy in Glaucoma Diagnosis and Treatment.Journal of clinical medicine · 2025Review
- From images to health insight: integrating MLLM, NLP, and objective Q-sorting of nursing-home built environment orientations.Frontiers in medicine · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
Abstract
Purpose: The aim is to assess GPT-4V's (OpenAI) diagnostic accuracy and its capability to identify glaucoma-related features compared to expert evaluations. Design: Evaluation of multimodal large language models for reviewing fundus images in glaucoma. Subjects: A total of 300 fundus images from 3 public datasets (ACRIMA, ORIGA, and RIM-One v3) that included 139 glaucomatous and 161 nonglaucomatous cases were analyzed. Methods: Preprocessing ensured each image was centered on the optic disc. GPT-4's vision-preview model (GPT-4V) assessed each image for various glaucoma-related criteria: image quality, image gradability, cup-to-disc ratio, peripapillary atrophy, disc hemorrhages, rim thinning (by quadrant and clock hour), glaucoma status, and estimated probability of glaucoma. Each image was analyzed twice by GPT-4V to evaluate consistency in its predictions. Two expert graders independently evaluated the same images using identical criteria. Comparisons between GPT-4V's assessments, expert evaluations, and dataset labels were made to determine accuracy, sensitivity, specificity, and Cohen kappa. Main Outcome Measures: The main parameters measured were the accuracy, sensitivity, specificity, and Cohen kappa of GPT-4V in detecting glaucoma compared with expert evaluations. Results: GPT-4V successfully provided glaucoma assessments for all 300 fundus images across the datasets, although approximately 35% required multiple prompt submissions. GPT-4V's overall accuracy in glaucoma detection was slightly lower (0.68, 0.70, and 0.81, respectively) than that of expert graders (0.78, 0.80, and 0.88, for expert grader 1 and 0.72, 0.78, and 0.87, for expert grader 2, respectively), across the ACRIMA, ORIGA, and RIM-ONE datasets. In Glaucoma detection, GPT-4V showed variable agreement by dataset and expert graders, with Cohen kappa values ranging from 0.08 to 0.72. In terms of feature detection, GPT-4V demonstrated high consistency (repeatability) in image gradability, with an agreement accuracy of ≥89% and substantial agreement in rim thinning and cup-to-disc ratio assessments, although kappas were generally lower than expert-to-expert agreement. Conclusions: GPT-4V shows promise as a tool in glaucoma screening and detection through fundus image analysis, demonstrating generally high agreement with expert evaluations of key diagnostic features, although agreement did vary substantially across datasets. Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.