Evidence map›Paper›PMID 39918656›Full record

ArticleInternational ophthalmology2025

Artificial intelligence with ChatGPT 4: a large language model in support of ocular oncology cases.

Federico Giannuzzi, Matteo Mario Carlà, Lorenzo Hu, Valentina Cestrone, Carmela Grazia Caputo, Maria Grazia Sammarco, Gustavo Savino, Stanislao Rizzo, Maria Antonietta Blasi, Monica Maria Pagliara

Abstract read
PubMed Publisher
In one paragraph

Article in International ophthalmology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Review
  2. Review
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Federico Giannuzzi *Ophthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy. federico.giannuzzi@gmail.com.
Matteo Mario Carlà *Ophthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Lorenzo HuOphthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Valentina CestroneOphthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Carmela Grazia CaputoOphthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Maria Grazia SammarcoOcular Oncology Unit, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Gustavo SavinoOcular Oncology Unit, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Stanislao RizzoOphthalmology Department, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Maria Antonietta BlasiOcular Oncology Unit, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.
Monica Maria PagliaraOcular Oncology Unit, "Fondazione Policlinico Universitario A. Gemelli, IRCCS", 00168, Rome, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeTo evaluate ChatGPT's ability to analyze comprehensive case descriptions of patients with uveal melanoma and provide recommendations for the most appropriate management.

designRetrospective analysis of ocular oncology patients' medical records. SUBJECTS: Forty patients treated for uveal melanoma between May 2019 and October 2023.

methodsWe uploaded each case description into the ChatGPT interface (version 4.0) and asked the model to provide realistic treatment options by asking the question, "What type of treatment do you recommend?" The accuracy of decisions produced by ChatGPT was compared to those recorded in patients' files and the treatment recommendations provided by three ocular oncologists, each with more than 10 years of experience.

main outcome measuresThe primary objective of this research was to assess the accuracy of ChatGPT replies in ocular oncology cases, analyzing its competence in both straightforward and intricate situations. Our secondary purpose was to assess the concordance between the responses of ChatGPT and those of ocular oncology specialists when faced with analogous clinical scenarios.

resultsChatGPT's surgical choices matched those in patients' files in 55% of cases (22 out of 40). ChatGPT options were agreed upon by 50%, 55%, and 57% of the three ocular oncology specialists. The investigation revealed significant differences between ChatGPT's responses and those of the three cancer specialists when compared to patients' files (p = 0.003, p = 0.001, and p = 0.001). ChatGPT's surgical responses matched with patient data in 18 out of 24 cases (75%), excluding enucleation cases. The decisions matched with the three ocular oncology specialists in 17/24, 18/24, and 18/24 cases, reflecting agreements of 70%, 75%, and 75%, respectively. The decisions made by ChatGPT were not significantly different from those of the three professionals in this cohort (p = 0.50, p = 0.36, and p = 0.36 for ChatGPT compared to specialists 1, 2, and 3).

conclusionChatGPT exhibited a level of proficiency that was comparable to that of trained ocular oncology specialists. However, it exhibited its limitations when evaluating more complex scenarios, such as extrascleral extension or infiltration of the optic nerve, when a comprehensive evaluation of the patient is therefore necessary.

Indexed as

Artificial IntelligenceMelanomaUveal NeoplasmsAdultAgedFemaleGenerative Artificial IntelligenceHumansLarge Language ModelsMaleMedical OncologyMiddle AgedRetrospective StudiesUveal MelanomaArtificial intelligence (AI)ChatGPTLarge language models (LLM)Ocular oncology

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.