ArticleJournal of endocrinological investigation2026
Artificial intelligence in endocrine practice: comparing ChatGPT, Gemini, and Claude for adrenal incidentaloma care.
Article in Journal of endocrinological investigation, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
- Diagnostic Performance and Workup Efficiency of Large Language Models in Secondary Hypertension: A Blinded Comparative Study.Diagnostics (Basel, Switzerland) · 2026Article
- Automated identification of incidentalomas requiring follow-up: A multi-anatomy evaluation of LLM-based and supervised approaches.Journal of biomedical informatics · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
purposeThe clinical use of artificial intelligence (AI) is expanding in endocrinology, yet the performance of large language models (LLMs) in managing adrenal incidentalomas remains uncertain. To compare the performance of four LLMs-ChatGPT-4o, ChatGPT-o1, Google Gemini 2.0, and Claude 3.5-on guideline-based queries and clinical scenarios involving adrenal incidentalomas.
methodsIn this cross-sectional study, 34 guideline-derived questions and four case scenarios were presented to the LLMs, covering diagnosis, treatment and follow-up, patient questions, and clinical cases. Six endocrinologists evaluated responses using Likert scales assessing hallucination tendency, quality, usability, reliability, and accuracy. Readability metrics and word counts were also analyzed.
resultsNo significant differences were found between models in diagnosis (p = 0.86-0.72), treatment and follow-up (p = 0.46-0.10), and patient question (p = 0.78-0.10) categories. However, in complex cases, ChatGPT-4o outperformed ChatGPT-o1 with higher scores in hallucination control (6.5 ± 0.8 vs. 4.8 ± 0.8), quality (6.2 ± 0.8 vs. 5.0 ± 0.6), and usability (4.5 ± 0.8 vs. 3.3 ± 0.5) (all p < 0.05). Readability analysis revealed high text complexity (Flesch-Kincaid Grade Level: 10.6-17.4), and inter-rater reliability was excellent (intraclass correlation coefficient: 0.876-0.961, p < 0.001).
conclusionLLMs show potential as decision-support tools in adrenal incidentaloma management. While their performance is comparable in routine tasks, significant differences arise in complex cases, highlighting the need for model selection, human oversight, and attention to readability in endocrine practice.
Indexed as
Identifiers
41055652What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.