Evidence map›Paper›PMID 42719386›Full record

ArticleDigital health

Dynamic alignment of large language models for evidence-grounded heart failure decision support.

Lu Liu, Chenchen Dong, Yunbo Ba, Haihong Yan, Xiaoxiao Tang, Yu Sun, Huilin Chen, Boyuan Shi, Qin Yu, Shulong Zhang

Abstract read
In one paragraph

Article in Digital health. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Lu LiuDalian Medical University, Dalian, Liaoning, China.ORCID https://orcid.org/0009-0002-1926-0470
Chenchen DongHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.
Yunbo BaHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.
Haihong YanHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.
Xiaoxiao TangHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.
Yu SunHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.
Huilin ChenFaculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia.
Boyuan ShiSchool of Biomedical Engineering, Tsinghua University, Beijing, China.ORCID https://orcid.org/0009-0000-7210-8086
Qin YuDalian Medical University, Dalian, Liaoning, China.
Shulong ZhangHeart Centre, Affiliated Zhongshan Hospital of Dalian University, Dalian, Liaoning, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objectives: Large language models (LLMs) are increasingly studied for clinical decision support, but high-risk cardiology exposes persistent weaknesses in hallucination control, guideline adherence, and medication-safety reasoning. Heart failure with reduced ejection fraction (HFrEF) is a demanding test case because safe care requires structured guideline-directed therapy, comorbidity-aware monitoring, and reliable risk warnings. Methods: We developed a dynamic alignment framework using 1087 retrospective HFrEF cases from Affiliated Zhongshan Hospital of Dalian University. An open-source LLaMA-3.1 backbone was optimized through four sequential stages: continual pre-training for heart-failure domain adaptation, supervised fine-tuning for structured clinical responses, reinforcement policy optimization for safety-oriented alignment, and retrieval-augmented generation for guideline grounding. Models were assessed with dual-track clinical and linguistic metrics. Results: LLaMA-3.1 was the strongest supervised baseline, but supervised fine-tuning alone did not fully resolve guideline-adherence limitations. Staged alignment produced a measurable Alignment Tax: the final retrieval-grounded variant improved the Clinical Score from 0.716 to 0.864 and reached a Guideline Score of 0.881, while BLEU-4 decreased from 0.371 to 0.272. The decline in surface overlap coincided with stronger risk safety, stricter structure, and more guideline-directed outputs. Conclusions: Dynamic alignment shifted the model from linguistic mimicry toward clinically constrained HFrEF decision support. These findings suggest that staged optimization with policy alignment and retrieval grounding can improve evidence-based recommendations, while conventional language-overlap metrics may underestimate clinically safer generation.

Indexed as

clinical decision support systemsguideline-directed medical therapyheart failure with reduced ejection fractionlarge language modelsretrieval-augmented generation

Identifiers

PMID42719386
PMCPMC13554667

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.