Evidence map›Paper›PMID 41695391›Full record

ArticleFrontiers in cell and developmental biology2026

Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering.

Ping Zhang, Jiaoman Wang, Xinya Hu, Xiaoqing Wang, Xianming Fan, Wei Chi, Weihua Yang

Abstract read
In one paragraph

Article in Frontiers in cell and developmental biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Ping Zhang *Department of Ophthalmology, Shenzhen People's Hospital, The Second Clinical Medical College of Jinan University, Shenzhen, Guangdong, China.
Jiaoman Wang *Eye Hospital and School of Ophthalmology and Optometry, Wenzhou Medical University, Wenzhou, Zhejiang, China.
Xinya HuShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Xiaoqing WangShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Xianming FanShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Wei ChiShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.
Weihua YangShenzhen Eye Hospital, Shenzhen Eye Medical Center, Southern Medical University, Shenzhen, Guangdong, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The application of large language models (LLMs) in medicine is rapidly advancing, showing particular promise in specialized fields like ophthalmology. However, existing research has predominantly focused on validating individual models, with a notable scarcity of systematic comparisons between multiple state-of-the-art LLMs. Objective: To systematically evaluate the performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 on ophthalmology question-answering tasks, with a specific focus on response consistency and factual accuracy. Methods: A total of 300 single-best-answer multiple-choice questions were sampled from the StatPearls ophthalmology question bank. The questions were categorized into four difficulty levels (Levels 1-4) based on the inherent difficulty ratings provided by the database. Each model provided independent answers three times under two distinct prompting strategies: a direct neutral prompt and a role-based prompt. Fleiss' kappa (κ) was used to assess inter-run response consistency, and overall accuracy was employed as the primary performance metric. Results: Accuracy: Gemini-3-Flash achieved the highest overall accuracy (83.3%), followed by GPT-o3 (79.2%) and DeepSeek-R1 (74.4%). GPT-4 (69.9%) and GPT-5 (69.1%) demonstrated the lowest accuracies. Consistency: GPT-o3 demonstrated the highest decision stability (κ = 0.966), followed by DeepSeek-R1 (κ = 0.904) and Gemini-3-Flash (κ = 0.860). GPT-5 exhibited the lowest stability (κ = 0.668). Influencing Factors: Prompting strategies did not significantly affect model accuracy. While Gemini-3-Flash remained stable across difficulty levels, DeepSeek-R1 and GPT-o3 showed enhanced relative performance on more complex tasks. Conclusion: GPT-o3 and Gemini-3-Flash achieve superior stability and accuracy in ophthalmology Question Answering (QA), making them suitable for high-stakes clinical decision support. The open-source model DeepSeek-R1 shows competitive potential, especially in complex tasks. Notably, GPT-5 failed to surpass its predecessor in both accuracy and consistency in this specialized domain. Prompt engineering has a limited impact on performance for closed-ended medical questions. Future work should extend to multimodal integration and real-world clinical validation to enhance the practical utility and reliability of LLMs in medicine.

Indexed as

artificial intelligence (AI)clinical decision supportlarge language model (LLM)medical educationophthalmology

Identifiers

PMID41695391
PMCPMC12894337

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.