Evidence map›Paper›PMID 37980501›Full record

ArticleInternational journal of retina and vitreous2023

"Application and accuracy of artificial intelligence-derived large language models in patients with age related macular degeneration".

Lorenzo Ferro Desideri, Janice Roth, Martin Zinkernagel, Rodrigo Anguita

Abstract read
In one paragraph

Article in International journal of retina and vitreous, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 25 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
25citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

25 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Trial
  3. Trial
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Review
  16. Article
  17. Article
  18. Article
  19. The digital age in retinal practice.International journal of retina and vitreous · 2024
    Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Lorenzo Ferro DesideriDepartment of Ophthalmology, Inselspital, University Hospital of Bern, Bern, Switzerland. lorenzoferrodes@gmail.com.
Janice RothDepartment of Ophthalmology, Inselspital, University Hospital of Bern, Bern, Switzerland.
Martin ZinkernagelDepartment of Ophthalmology, Inselspital, University Hospital of Bern, Bern, Switzerland.
Rodrigo AnguitaDepartment of Ophthalmology, Inselspital, University Hospital of Bern, Bern, Switzerland.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

introductionAge-related macular degeneration (AMD) affects millions of people globally, leading to a surge in online research of putative diagnoses, causing potential misinformation and anxiety in patients and their parents. This study explores the efficacy of artificial intelligence-derived large language models (LLMs) like in addressing AMD patients' questions.

methodsChatGPT 3.5 (2023), Bing AI (2023), and Google Bard (2023) were adopted as LLMs. Patients' questions were subdivided in two question categories, (a) general medical advice and (b) pre- and post-intravitreal injection advice and classified as (1) accurate and sufficient (2) partially accurate but sufficient and (3) inaccurate and not sufficient. Non-parametric test has been done to compare the means between the 3 LLMs scores and also an analysis of variance and reliability tests were performed among the 3 groups.

resultsIn category a) of questions, the average score was 1.20 (± 0.41) with ChatGPT 3.5, 1.60 (± 0.63) with Bing AI and 1.60 (± 0.73) with Google Bard, showing no significant differences among the 3 groups (p = 0.129). The average score in category b was 1.07 (± 0.27) with ChatGPT 3.5, 1.69 (± 0.63) with Bing AI and 1.38 (± 0.63) with Google Bard, showing a significant difference among the 3 groups (p = 0.0042). Reliability statistics showed Chronbach's α of 0.237 (range 0.448, 0.096-0.544).

conclusionChatGPT 3.5 consistently offered the most accurate and satisfactory responses, particularly with technical queries. While LLMs displayed promise in providing precise information about AMD; however, further improvements are needed especially in more technical questions.

Indexed as

Artificial IntelligenceArtificial intelligence in ophthalmologyDry macular degenerationLarge language modelsLLMsMacular edemaWet macular degeneration

Identifiers

PMID37980501
PMCPMC10657493

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.