Evidence map›Paper›PMID 42249916›Full record

ArticleUrolithiasis2026

Performance of large language models in urological decision support: a guideline-based comparative evaluation in urolithiasis.

Adem Tunçekin, Yasin Aktaş

Abstract readComparative Study
In one paragraph

Article in Urolithiasis, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Adem TunçekinFaculty of Medicine, Department of Urology, Uşak University, Uşak, Turkey. dr_adem65@hotmail.com.ORCID http://orcid.org/0000-0001-9951-2556
Yasin AktaşFaculty of Medicine, Department of Urology, Uşak University, Uşak, Turkey.ORCID http://orcid.org/0000-0001-5255-3780

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Large language models (LLMs) are increasingly investigated for their potential role in guideline-based clinical information support. However, their consistency with subspecialty guidelines, particularly in urolithiasis, remains underexplored. This study aimed to evaluate the performance of four large language models (LLMs); GPT-4, GPT-4-turbo, Claude, and Gemini in generating guideline-concordant responses to clinical questions related to urolithiasis. A total of 105 clinical questions were independently developed by the authors based on urolithiasis management principles. Each LLM generated responses in two separate sessions. Two experienced urologists evaluated the outputs for accuracy and concordance with guideline recommendations. Inter-rater agreement analysis demonstrated fair agreement between evaluators. Differences across models and guideline categories were assessed using appropriate statistical tests. All four LLMs demonstrated high guideline adherence, with mean total scores ranging from 86.5 ± 5.2 (Gemini) to 92.8 ± 3.1 (Claude). Claude achieved the highest correlation with expert ratings (r = 0.94, p < 0.01). There were no statistically significant differences across models or among the nine clinical categories (p > 0.05). Session-to-session repeatability was also high for all models, with intra-model correlation coefficients exceeding 0.90. LLMs, particularly Claude, can provide reliable, guideline-consistent answers to urolithiasis-related clinical queries. Their consistent performance across themes suggests utility as adjunctive informational tools for guideline-based urological education and support, although further validation in real-world clinical settings remains necessary.

Indexed as

Decision Support Systems, ClinicalLarge Language ModelsUrolithiasisUrologyGuideline AdherenceHumansPractice Guidelines as TopicArtificial intelligenceClinical decision supportGuideline adherenceLarge language modelsNatural language processingUrolithiasis

Identifiers

PMID42249916
PMCPMC13242490

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.