Evidence map›Paper›PMID 42746177›Full record

ArticleFrontiers in psychiatry2026

HMDF-Net: a transfer learning-based heterogeneous multimodal dynamic fusion network for depression detection among inmates in correctional facilities.

Zhifei Xu, Shiyun Shao, Jiagang Dong, Ziguan Wei, Yichao Zhang

Abstract read
In one paragraph

Article in Frontiers in psychiatry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Zhifei XuZhejiang Provincial Academy of Judicial Administration, Zhejiang Police Vocational Academy, Hangzhou, China.
Shiyun ShaoZhejiang Provincial Engineering Research Center for Brain Cognition, Disease and Digital Medical Devices, School of Information Engineering, Hangzhou Medical College, Hangzhou, China.
Jiagang DongDepartment of Digital Technology, Zhejiang Yuying College of Vocational Technology, Hangzhou, China.
Ziguan WeiZhejiang Provincial Academy of Judicial Administration, Zhejiang Police Vocational Academy, Hangzhou, China.
Yichao ZhangZhejiang Provincial Engineering Research Center for Brain Cognition, Disease and Digital Medical Devices, School of Information Engineering, Hangzhou Medical College, Hangzhou, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: In correctional institutions, inmates frequently suffer from mental health issues (such as depression), which can easily lead to emergencies like suicide, self-harm, and violent conflicts, posing severe challenges to security management and recidivism prevention. Traditional assessment methods have limitations such as subjectivity, lag, and low efficiency, which may result in misjudgment, omission, and delayed judgment. To address these issues, this study aims to develop an automated depression screening model suitable for the complex scenarios in correctional institutions to achieve objective, accurate, dynamic, and efficient non-contact mental health risk monitoring and early warning. Methods: A Heterogeneous Multimodal Dynamic Fusion Network (HMDF-Net) is proposed. This method first constructs a multimodal dataset for inmates and adopts a transfer learning strategy to alleviate the problem of data scarcity. HMDF-Net integrates VGG19, wav2vec 2.0, and BERT to extract visual (facial micro-expressions), audio (speech prosody), and text (dialogue semantics) features respectively. A temporal attention dynamic fusion mechanism is introduced to dynamically assign weights according to the real-time signal quality of each modal channel. Results: HMDF-Net achieves an accuracy of 87.5% and a recall rate of 100% on the test set, significantly outperforming various comparison methods. The analysis of the dynamic attention mechanism reveals an intelligent hierarchical fusion paradigm of "visual as the main (48%), text as the auxiliary (32%), and audio as the supplement (20%)", verifying the core role of the visual modality in combating emotional concealment and the decision-making logic of the multimodal dynamic fusion model in complex environments. Conclusion: We present HMDF-Net, a cross-modal fusion network equipped with temporal attention that adjusts feature weights according to signal quality to suppress low-quality multimodal inputs. This framework offers a feasible option for scalable non-invasive mental health screening within correctional settings. Our observations suggest dataset limitations unique to this field matter more to model performance than available computing resources.

Indexed as

correctional mental healthdepression detectiondynamic attention mechanismfew-shot learningmultimodal fusionnon-cooperative environmenttransfer learning

Identifiers

PMID42746177
PMCPMC13576147

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.