Evidence map›Paper›PMID 41781977›Full record

ArticleBioData mining2026

Disease- and gene-specific deep learning for pathogenicity prediction of rare missense variants in cancer predisposition genes.

Da-Bin Lee, Hyun-Uk Kang, Kyu-Baek Hwang

Abstract read
In one paragraph

Article in BioData mining, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Da-Bin LeeDepartment of Computer Science and Engineering, Graduate School, Soongsil University, Seoul, 06978, Korea.
Hyun-Uk KangDepartment of Computer Science and Engineering, Graduate School, Soongsil University, Seoul, 06978, Korea.
Kyu-Baek HwangDepartment of Computer Science and Engineering, Graduate School, Soongsil University, Seoul, 06978, Korea. kbhwang@ssu.ac.kr.

Funding

the National Research Foundation of Korea NRF2022R1F1A1072718
6 · The paper itself

Abstract

backgroundHereditary cancers frequently arise from germline pathogenic variants, yet only a small proportion of reported variants have been clinically classified, leaving most missense variants unresolved as variants of uncertain significance (VUS). Although recent machine-learning approaches have explored disease-specific or gene-specific contexts to improve pathogenicity prediction, these models remain fundamentally limited by the scarcity of labeled data and the underutilization of abundant VUS.

resultsWe propose a deep-learning framework that integrates autoencoder pretraining with a deep ensemble strategy to improve variant pathogenicity prediction, effectively leverage unlabeled VUS during pretraining, and reduce uncertainty arising from limited training samples. To validate each component of our framework, we evaluated its performance under both disease-specific and gene-specific training setups. Experiments on ClinVar variants from BRCA1, BRCA2, MLH1, and MSH2 showed that our framework achieved the best performance in the gene-specific setup for BRCA1—likely because BRCA1 contains substantially more gene-specific training data than the other genes—whereas the disease-specific setup yielded superior results for the remaining genes, which had comparatively limited gene-specific samples. Overall, our method significantly outperformed existing approaches. We also introduce an interpretability approach that provides variant-level importance profiles across pathogenicity classes, thereby enhancing transparency and clinical applicability. Moreover, by projecting feature-level importance scores into a two-dimensional space, we demonstrate that pretraining enables the model to learn distinctly different feature representations, illustrating how pretraining and ensemble learning synergistically contribute to improved predictive performance.

conclusionsOur framework preserves the specificity of disease- and gene-specific approaches, overcomes data scarcity through VUS-guided pretraining and ensembling, and offers interpretable outcomes that may be helpful for clinical decision support. Moreover, our results suggest a promising direction for pathogenicity prediction of rare missense variants and indicate that the proposed framework may be extendable to additional genes under appropriate data and modeling conditions.

Indexed as

Deep ensembleDeep neural networksDenoising autoencodersPredisposition cancer genesRare missense variantRepresentation learningVariant pathogenicity prediction

Identifiers

PMID41781977
PMCPMC12964711

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.