Evidence map›Paper›PMID 42819102›Full record

ArticleFrontiers in endocrinology2026

From risk classification to clinical action: a prespecified paired pilot benchmark of public large language model interfaces for diabetes-related foot ulcer prevention.

Yang Wen, Liyuan Chen, Xin Deng, Jiaping Lan, Xin Huang, Yang Liu, Lin Chen, Lei Li

Abstract read
In one paragraph

Article in Frontiers in endocrinology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Yang Wen *Department of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Liyuan Chen *Medical Department, Suining Central Hospital, Suining, Sichuan, China.
Xin DengDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Jiaping LanDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Xin HuangDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Yang LiuDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Lin ChenDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.
Lei LiDepartment of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, Sichuan, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) may assist in the prevention of diabetes-related foot ulcers; however, their performance in classification may not translate effectively to context-dependent decisions. Objective: This study aimed to evaluate the accuracy, clinical actionability, reproducibility, and safety of five public LLM web interfaces in the context of the International Working Group on the Diabetic Foot (IWGDF) 2023 risk stratification and preventive management. Methods: A prespecified, paired, blinded, noninterventional pilot benchmark was conducted using 20 translated, de-identified, curated cases structured in a fixed clinical sequence and evenly distributed across the four IWGDF risk categories. These input conditions represent an idealized, best-case benchmark rather than routine clinical documentation. Each interface assessed all cases under standardized no-search conditions. The primary outcome was exact agreement with a frozen multidisciplinary consensus risk category. Additional outcomes included screening frequency, need for referral, specialty and urgency of referrals, completeness, interreviewer reliability, repeated-generation stability, and safety. Results: Each interface classified 20/20 cases in concordance with the frozen IWGDF reference (100%; Wilson 95% CI 83.9%-100.0%). The 16.1-percentage-point interval below the observed ceiling indicates limited precision and remains compatible with clinically meaningful error in new cases. Of 100 primary outputs, agreement was observed in 89/100 (89.0%) for referral need, 68/75 (90.7%) for referral specialty, and 92/100 (92.0%) for urgency. Median information-item and mandatory-measure coverage was 100.0% for both measures; however, completeness scoring had limited inter-reviewer reliability and should be interpreted cautiously. In an exploratory, hypothesis-generating four-case repeated-generation substudy, all-three-generation consensus concordance was observed in 16/20 interface-case combinations for referral need, 14/15 eligible combinations for specialty, and 19/20 combinations for urgency. These descriptive counts are not estimates of failure probability or tail behavior. Under the prespecified curated no-search benchmark conditions, no major safety errors were observed among 140 outputs; this absence of observed events does not establish safety in routine clinical use. Conclusions: Performance was highest for structured guideline mapping, though reliability diminished in referral and individualized management across repeated generations. These findings highlight the necessity for auditable, clinician-supervised decision support instead of autonomous deployment.

Indexed as

BenchmarkingDiabetic FootLarge Language ModelsHumansPilot ProjectsReproducibility of ResultsRisk Assessmentartificial intelligenceclinical decision supportdiabetic footlarge language modelspreventive carerisk stratification

Identifiers

PMID42819102
PMCPMC13623619

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.