Evidence map›Paper›PMID 42421754›Full record

ArticleOphthalmology science2026

Image-Quality-Aware Multimodal Artificial Intelligence for Automated Structured OCT Report Generation in Glaucoma Evaluation.

Jalil Jalili, Yashraj Gavhane, Evan Walker, Anna Heinke, Christopher Bowd, Akram Belghith, Massimo A Fazio, Christopher A Girkin, C Gustavo De Moraes, Jeffrey M Liebmann and 4 more

Abstract read
In one paragraph

Article in Ophthalmology science, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Jalil JaliliDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Yashraj GavhaneDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Evan WalkerDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Anna HeinkeDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Christopher BowdDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Akram BelghithDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Massimo A FazioDepartment of Ophthalmology and Vision Sciences, University of Alabama at Birmingham, Birmingham, Alabama.
Christopher A GirkinDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
C Gustavo De MoraesDepartment of Ophthalmology, Harkness Eye Institute, Bernard and Shirlee Brown Glaucoma Research Laboratory, New York, New York.
Jeffrey M LiebmannDepartment of Ophthalmology, Harkness Eye Institute, Bernard and Shirlee Brown Glaucoma Research Laboratory, New York, New York.
Sally L BaxterDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Robert N WeinrebDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Linda M ZangwillDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Mark ChristopherDivision of Ophthalmology Informatics and Data Science, Viterbi Family Department of Ophthalmology, Shiley Eye Institute, University of California San Diego, La Jolla, California.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: To develop an explainable multimodal large language model (MM-LLM) that (1) screens optic nerve head (ONH) OCT circle scans for quality and (2) generates structured clinical reports that include glaucoma diagnosis and sector-wise retinal nerve fiber layer (RNFL) thinning assessments. Design: A retrospective cohort study using longitudinal data from the Diagnostic Innovations in Glaucoma Study and the African Descent and Glaucoma Evaluation Study. Participants: A total of 43 849 Spectralis circumpapillary B-scans centered on the ONH from 1310 subjects, including 1331 glaucomatous and 867 healthy eyes. Methods: An MM-LLM (Llama 3.2 Vision-Instruct model) was fine-tuned to generate clinical descriptions of OCT imaging data. Training data included paired OCT images and automatically generated, structured clinical reports that described global and sectoral RNFL thinning. Poor-quality scans were labeled as unusable and paired with a fixed refusal statement. The model was evaluated on a held-out test set for 3 tasks: quality assessment, glaucoma detection, and RNFL thinning classification across 7 anatomical sectors. Evaluation metrics included accuracy, sensitivity, specificity, precision, and F1-score. Model description quality was also evaluated using standard text evaluation metrics (BLEU, ROUGE, METEOR, and BERTScore). Main Outcome Measures: Diagnostic accuracy metrics for each task; text evaluation metrics for description quality. Results: The model achieved 0.90 accuracy and 0.98 specificity for quality triage. For glaucoma detection, accuracy was 0.86 (sensitivity 0.93, specificity 0.65, and F1-score 0.91). Retinal nerve fiber layer thinning prediction accuracy ranged from 0.83 to 0.94, with the highest performance in global, temporal, temporal superior, and temporal inferior sectors. Text generation scores (mean ± standard deviation) showed strong alignment with reference reports (BLEU: 0.82 ± 0.19; ROUGE-1: 0.94 ± 0.08; ROUGE-2: 0.87 ± 0.17; ROUGE-L: 0.92 ± 0.11; BERTScore-F1: 0.99 ± 0.02). Stratified analysis revealed better RNFL thinning detection in moderate-to-advanced glaucoma cases, especially in temporal sectors, while performance in nasal regions was better for mild cases. Conclusions: The fine-tuned MM-LLM generated accurate clinical descriptions based on OCT imaging. The model achieved high accuracy in identifying image quality issues and detecting glaucoma. The model provided sectoral descriptions of RNFL thinning to support clinical OCT evaluation. This approach shows potential as a scalable tool for clinical decision support, but further validation across additional datasets is needed. Financial Disclosures: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Indexed as

Clinical report generationGlaucoma detectionMultimodal large language modelQuality triageRetinal nerve fiber layer

Identifiers

PMID42421754
PMCPMC13343140

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.