Evidence map›Paper›PMID 42174361›Full record

ArticleHepatology international2026

Machine learning outperforms serum creatinine for early risk stratification of hepatorenal syndrome: a prospective single-center validation.

Yuli Song, Weixiao Shi, Chengchen Yang, Xiaochen Yang, Chengbo Yu

Abstract read
PubMed Publisher
In one paragraph

Article in Hepatology international, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Yuli SongState Key Laboratory for Diagnosis and Treatment of Infectious DiseasesCollaborative Innovation Center for Diagnosis and Treatment of Infectious DiseasesSchool of Medicine, National Clinical Research Center for Infectious DiseasesThe First Affiliated HospitalZhejiang University, 79 Qingchun Road, Hangzhou, 310003, Zhejiang, People's Republic of China.
Weixiao ShiState Key Laboratory for Diagnosis and Treatment of Infectious DiseasesCollaborative Innovation Center for Diagnosis and Treatment of Infectious DiseasesSchool of Medicine, National Clinical Research Center for Infectious DiseasesThe First Affiliated HospitalZhejiang University, 79 Qingchun Road, Hangzhou, 310003, Zhejiang, People's Republic of China.
Chengchen YangState Key Laboratory for Diagnosis and Treatment of Infectious DiseasesCollaborative Innovation Center for Diagnosis and Treatment of Infectious DiseasesSchool of Medicine, National Clinical Research Center for Infectious DiseasesThe First Affiliated HospitalZhejiang University, 79 Qingchun Road, Hangzhou, 310003, Zhejiang, People's Republic of China.
Xiaochen YangZhejiang Provincial Tongde Hospital, Hangzhou, Zhejiang, People's Republic of China.
Chengbo YuState Key Laboratory for Diagnosis and Treatment of Infectious DiseasesCollaborative Innovation Center for Diagnosis and Treatment of Infectious DiseasesSchool of Medicine, National Clinical Research Center for Infectious DiseasesThe First Affiliated HospitalZhejiang University, 79 Qingchun Road, Hangzhou, 310003, Zhejiang, People's Republic of China. yuchengbo1974@zju.edu.cn.ORCID http://orcid.org/0000-0003-0766-8794

Funding

Innovative Research Group Project of the National Natural Science Foundation of China 82200673
6 · The paper itself

Abstract

backgroundThe early diagnosis of hepatorenal syndrome (HRS) is constrained by the reliance on serum creatinine, a biomarker with well-documented limitations in sensitivity and timeliness, contributing to diagnostic delays and adverse outcomes. Machine learning (ML) offers a potential solution, but its translation is hindered by two critical gaps: the absence of prospective, head-to-head validation against the clinical standard, and inadequate assessment of model robustness against temporal data distribution shift-a pivotal challenge for real-world deployment.

methodsWe developed an XGBoost model using a retrospective cohort of patients with decompensated cirrhosis (n = 464). Its performance was then prospectively validated in a completely independent, temporally distinct cohort (n = 269) in a direct comparison against the serum creatinine diagnostic standard. The prospective cohort was further split into sequential subsets to assess short-term internal temporal robustness. The evaluation framework encompassed diagnostic performance, interpretability (SHAP analysis), clinical utility (decision curve analysis), and health economic evaluation.

resultsThe model demonstrated exceptional discriminatory accuracy, with an area under the receiver operating characteristic curve (AUC) of 0.994 (95% bias-corrected and accelerated bootstrap CI: 0.985-0.998). This represented a substantial and statistically significant improvement over serum creatinine-based diagnosis (AUC = 0.803; p < 0.001). The model's predictions were well-calibrated, and its performance remained stable in the limited internal temporal robustness assessment. The health economic evaluation established the model's decisive cost-effectiveness, with an incremental cost-effectiveness ratio (ICER) of 571 EUR per quality-adjusted life year (QALY), substantially below major international willingness-to-pay thresholds, including Indonesia's benchmark of 3853 EUR/QALY.

conclusionThis study provides three key contributions: (1) conclusive, prospective evidence that an ML model significantly outperforms the current serum creatinine standard for HRS risk stratification; (2) a novel methodological framework for proactively assessing temporal robustness, which demonstrated stability in a short-term, single-center setting; and (3) compelling evidence of the model's accuracy, clinical utility, and cost-effectiveness within a rigorous, single-center prospective validation. While these results establish a high-fidelity proof of concept, the essential next steps are external validation across multiple centers to confirm generalizability and the assessment of resilience to longer-term, real-world dataset shifts before any consideration of broader clinical implementation.

Indexed as

Cost-effectivenessHepatorenal syndromeMachine learningProspective validationRisk stratificationTemporal distribution shift

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.