Evidence map›Paper›PMID 41885289›Full record

ArticleEuropean thyroid journal2026

Promise and pitfalls of AI chatbots in complex decision-making for thyroid nodules and papillary thyroid cancer.

Grigoris Effraimidis, Athanasios Kasotas, Sofia Varsami, Eleni Sazakli, Olga Karapanou, Katerina Saltiki, Marina Michalaki

Abstract read
In one paragraph

Article in European thyroid journal, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Grigoris EffraimidisFaculty of Medicine, School of Health Sciences, University of Thessaly, Larissa, Greece.ORCID 0000-0001-7313-8391
Athanasios KasotasDepartment of Endocrinology and Metabolic Diseases, Larissa University Hospital, Larissa, Greece.
Sofia VarsamiBioinformatics, Faculty of Science, University of Copenhagen, Copenhagen, Denmark.
Eleni SazakliFaculty of Medicine, School of Health Science, University of Patras, Patras, Greece.
Olga KarapanouEndocrine Department, NIMTS Veteran's Hospital, Athens, Greece.ORCID 0000-0002-9978-5885
Katerina SaltikiEndocrine Unit, Department of Clinical Therapeutics, National and Kapodistrian University, Athens, Greece.
Marina MichalakiFaculty of Medicine, School of Health Science, University of Patras, Patras, Greece.ORCID 0000-0003-2073-8523

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Introduction: Artificial intelligence (AI) chatbots are increasingly used in medicine, but their reliability in scenarios with multiple management options is unclear. Indeterminate thyroid nodules and low- and low-to-intermediate-risk papillary thyroid carcinoma (PTC) represent such cases. Methods: In a nationwide web-based survey, 201 members of the Hellenic Endocrine Society evaluated 12 clinical vignettes on indeterminate thyroid nodules and low- and low-to-intermediate-risk PTC. Their responses were compared with those generated by four conversational AI models (ChatGPT, Gemini, Copilot, and DeepSeek) at two time points, 11 months apart. DeepSeek was assessed only at the second time point. Chatbot outputs were assessed for agreement with endocrinologists' predominant answers, concordance with the most guideline-consistent options (American and European Thyroid Association recommendations), temporal stability, and inter-model agreement. Results: Alignment between chatbots and endocrinologists' predominant responses was limited, reaching at most 25% across scenarios. In contrast, concordance with the most guideline-consistent options was higher, up to 83% (10/12 scenarios), depending on the model and time point. Across 12 scenarios, ChatGPT, Gemini, and Copilot changed their responses in 4, 7, and 5 scenarios, respectively, with some updates moving closer to, and others further from, guideline-based answers. Inter-model agreement ranged from 33 to 67%, indicating substantial variability among chatbots. Conclusion: AI chatbots show evolving but inconsistent performance in complex thyroid management scenarios. While guideline concordance can be relatively high, substantial variability across models, limited temporal reproducibility, and poor alignment with clinical practice highlight the need for ongoing longitudinal evaluation before safe integration into clinical decision-making.

Indexed as

Artificial IntelligenceClinical Decision-MakingThyroid Cancer, PapillaryThyroid NeoplasmsThyroid NoduleGenerative Artificial IntelligenceHumansIntelligent SystemsSurveys and Questionnairesartificial intelligencechatbotsclinical decision-makingpapillary thyroid cancersurveythyroid nodules

Identifiers

PMID41885289
PMCPMC13087872

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.