Evidence map›Paper›PMID 42704533›Full record

ArticleInternational ophthalmology2026

Generative artificial intelligence to augment ethical problem solving in ophthalmology: GPT-5.1 versus a human ethicist.

Daniel C Kelly, Ishan Chillikatil, Jonathon M Monroe, Rebika Khanal, Andrew Trippiedi, Chelsea-Jane Arcalas, Matthew R Claxton, Tochukwu Ndukwe, Elizabeth Pogrebniak, Joshua Barnett and 3 more

Abstract read
PubMed Publisher
In one paragraph

Article in International ophthalmology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Daniel C KellyDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Ishan ChillikatilDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Jonathon M MonroeDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Rebika KhanalDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Andrew TrippiediDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Chelsea-Jane ArcalasDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Matthew R ClaxtonDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Tochukwu NdukweDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Elizabeth PogrebniakDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Joshua BarnettDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Jacquelyn O'BanionDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Jeremy K JonesDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA.
Rebecca F NeusteinDepartment of Ophthalmology, Emory University School of Medicine, Atlanta, GA, USA. rebecca.neustein@emory.edu.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeWhile many large language models (LLMs) have been extensively investigated for their clinical decision-making capabilities, few studies have characterized their abilities to reason through complex, open-ended ethics cases. This study compared GPT-5.1 and expert human ethicist responses to real-world ethical scenarios specific to ophthalmology.

methodsTen ethical scenarios from the American Academy of Ophthalmology's Ask the Ethicist website were randomly selected and presented to GPT-5.1 using ChatGPT. AI-generated responses were subsequently compared to those of the expert human ethicist using conventional readability metrics. A panel of 10 physicians independently rated all responses via 6-point and 5-point Likert scales for both outcomes of likelihood-of-use in their own careers and perceived patient impact, respectively.

resultsGPT-5.1 versus human ethicist responses differed significantly on Flesch Reading Ease (10.9 ± 9.6 vs. 25.4 ± 10.9, p = 0.002) and Gunning Fog Index (21.3 ± 1.9 vs. 19.5 ± 2.9, p = 0.037). For likelihood-of-use, median ratings were 4.1 [3.8-4.3] for human ethicist versus 5.0 [4.7-5.2] for GPT-5.1 responses (p = 0.013). Median ratings for perceived patient impact of human ethicist versus GPT-5.1 responses were 3.6 [3.3-3.7] versus 3.9 [3.8-4.4], p = 0.008. Inter-rater reliability was moderate for both outcomes and response sources (ICC [2, 10]: 0.56-0.69).

conclusionsCompared to human ethicist responses, GPT-5.1 responses demonstrated lower readability; however, GPT-5.1 responses received significantly higher ratings for both outcomes of likelihood-of-use and perceived patient impact, respectively. These results advocate for further explorations of GPT-5.1 and other LLMs as potentially useful tools for evaluating common bioethical challenges in ophthalmology.

Indexed as

Clinical Decision-MakingEthicistsGenerative Artificial IntelligenceOphthalmologyProblem SolvingFemaleHumansLarge Language ModelsMaleArtificial intelligence (AI)Ethical reasoningEthics in ophthalmologyGPT-5.1Large language models (LLMs)

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.