Evidence map›Paper›PMID 41484266›Full record

ArticleGraefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie2026

Comparing the performance of four mainstream large language models on medical literature review generation: a human expert evaluation in SMILE surgery.

Mengyun Zhou, Fu Gui, Chong Ai, Ling Ling, Xian Zhang, Yalin Lu, Dongmei Han, Bin Zhao, Fei Zhong, Jie Liu and 6 more

Abstract readComparative Study
In one paragraph

Article in Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

16 authors.

Mengyun Zhou *Ophthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Fu Gui *Ophthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Chong Ai *Ophthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Ling LingNanchang Bright Eye Hospital, Nanchang, 330000, China.
Xian ZhangNanchang Bright Eye Hospital, Nanchang, 330000, China.
Yalin LuOphthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Dongmei HanShanghai Heping Eye Hospital, Shanghai, 200030, China.
Bin ZhaoShanghai Heping Eye Hospital, Shanghai, 200030, China.
Fei ZhongDepartment of Ophthalmology, First Affiliated Hospital of Gannan Medical University, Ganzhou, 341000, China.
Jie LiuXiaogan Central Hospital, Xiaogan, 432000, China.
Zeyu ZhuHeyou International Hospital, Shenzhen, 518000, China.
Jiayu LiThe First Affiliated Hospital of Xi'an Jiao Tong University, Xi'an, 710089, China.
Fei HuangOphthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Chuyang LinOphthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China.
Weifeng LiuOphthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China. 18970040725@163.com.
Jian XiongOphthalmic Center, The Second Affiliated Hospital, Jiangxi Medical College, Nanchang University, Nanchang, 330000, China. 894040417@qq.com.

Funding

National Natural Science Foundation of China 82260214Science and Technology Program of Jiangxi Provincial Health Commission 202210631
6 · The paper itself

Abstract

purposeTo systematically evaluate and compare the performance of four leading large language models (LLMs) in generating medical literature reviews across topics of varying research maturity, thereby providing insights for their effective and responsible application in academic writing.

methodsIn this comparative study, using standardized prompts, we instructed four leading LLMs (GPT-4, Gemini 2.5 Pro, Grok-3, and DeepSeek R1) to generate literature reviews on nine topics related to small incision lenticule extraction (SMILE) surgery. These topics were categorized into three groups by research maturity: well-researched, controversial, and open. Seven ophthalmology experts evaluated the generated content across four dimensions: quality, accuracy, bias, and relevance, while all references were verified for authenticity. Performance differences among models were evaluated using group comparison tests followed by post-hoc analysis.

resultsSignificant performance variations were identified across all four models and dimensions (p < 0.001). Specifically, Gemini ranked highest in content quality, accuracy, and bias control. In contrast, DeepSeek, despite its high-quality score, received the lowest relevance score. Grok-3 demonstrated the highest reference authenticity (p < 0.001), whereas GPT-4's was the lowest (p < 0.001). All models showed diminished performance on open topics and exhibited severe reference fabrication ("hallucinations").

conclusionRather than excelling universally, LLMs exhibit distinct and task-specific strengths that mandate a task-driven, hybrid strategy in tool selection. Reference fabrication was found to be a pervasive issue across all models, regardless of the task topic, elevating human verification from a best practice to an essential safeguard for academic integrity.

Indexed as

Corneal StromaCorneal Surgery, LaserLanguageMyopiaHumansLarge Language ModelsOphthalmologyAI evaluationLarge language modelsLiterature reviewMedical content generationSMILE surgery

Identifiers

PMID41484266
PMCPMC13091840

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.