Evidence map›Paper›PMID 41055652›Full record

ArticleJournal of endocrinological investigation2026

Artificial intelligence in endocrine practice: comparing ChatGPT, Gemini, and Claude for adrenal incidentaloma care.

Özge Baş Aksu, Rıfat Furkan Aydın, Asena Gökçay Canpolat, Özgür Demir, Mustafa Şahin, Rıfat Emral, Sevim Güllü

Abstract readComparative Study
PubMed Publisher
In one paragraph

Article in Journal of endocrinological investigation, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Özge Baş AksuDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey. ozgebasaksu@gmail.com.ORCID http://orcid.org/0000-0003-3124-9477
Rıfat Furkan AydınDepartment of Internal Medicine, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0003-0865-7835
Asena Gökçay CanpolatDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0003-1186-2960
Özgür DemirDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0001-6555-3579
Mustafa ŞahinDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0002-4718-0083
Rıfat EmralDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0002-5732-2284
Sevim GüllüDepartment of Endocrinology and Metabolism, Ankara University School of Medicine, Ankara, Turkey.ORCID http://orcid.org/0000-0002-0955-0717

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

purposeThe clinical use of artificial intelligence (AI) is expanding in endocrinology, yet the performance of large language models (LLMs) in managing adrenal incidentalomas remains uncertain. To compare the performance of four LLMs-ChatGPT-4o, ChatGPT-o1, Google Gemini 2.0, and Claude 3.5-on guideline-based queries and clinical scenarios involving adrenal incidentalomas.

methodsIn this cross-sectional study, 34 guideline-derived questions and four case scenarios were presented to the LLMs, covering diagnosis, treatment and follow-up, patient questions, and clinical cases. Six endocrinologists evaluated responses using Likert scales assessing hallucination tendency, quality, usability, reliability, and accuracy. Readability metrics and word counts were also analyzed.

resultsNo significant differences were found between models in diagnosis (p = 0.86-0.72), treatment and follow-up (p = 0.46-0.10), and patient question (p = 0.78-0.10) categories. However, in complex cases, ChatGPT-4o outperformed ChatGPT-o1 with higher scores in hallucination control (6.5 ± 0.8 vs. 4.8 ± 0.8), quality (6.2 ± 0.8 vs. 5.0 ± 0.6), and usability (4.5 ± 0.8 vs. 3.3 ± 0.5) (all p < 0.05). Readability analysis revealed high text complexity (Flesch-Kincaid Grade Level: 10.6-17.4), and inter-rater reliability was excellent (intraclass correlation coefficient: 0.876-0.961, p < 0.001).

conclusionLLMs show potential as decision-support tools in adrenal incidentaloma management. While their performance is comparable in routine tasks, significant differences arise in complex cases, highlighting the need for model selection, human oversight, and attention to readability in endocrine practice.

Indexed as

Adrenal Gland NeoplasmsArtificial IntelligenceEndocrinologyCross-Sectional StudiesFemaleGenerative Artificial IntelligenceHumansMaleAdrenal incidentalomaArtificial intelligenceChatGPTClaudeGemini

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.