ArticleBriefings in bioinformatics2024
ULDNA: integrating unsupervised multi-source language models with LSTM-attention network for high-accuracy protein-DNA binding site prediction.
Article in Briefings in bioinformatics, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 21 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
21 citing papers in PubMed, 1 synthesis or guideline pooled it.
- Advances in the Application of Protein Language Modeling for Nucleic Acid Protein Binding Site Prediction.Genes · 2024Pooled it
- DNAreader: accurate prediction of DNA-binding residues in structured and disordered proteins using transformers and contrastive learning.Nucleic acids research · 2026Article
- Protein-nucleic acid binding site prediction using interpretable Kolmogorov-Arnold networks with hypergraph representation learning.Bioinformatics (Oxford, England) · 2026Article
- Predicting protein-nucleic acid interactions via protein language models with biophysical and evolutionary priors.iScience · 2026Article
- TripleBind: a generalizable deep learning framework for protein-nucleic acid and protein-ligand binding sites prediction based on pre-trained protein language models.Molecular diversity · 2026Article
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations.Briefings in bioinformatics · 2026Article
- PRIMED: predicting DNA binding residues by leveraging pre-trained protein language models.Frontiers in artificial intelligence · 2026Article
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations.bioRxiv : the preprint server for biology · 2025Article
- Multimodal Cross-Attention Molecular Property Prediction for Text, Sequence, Graph, and Geometry.ACS omega · 2025Article
- M3Site: multiclass multimodal learning for protein active site identification and classification.Briefings in bioinformatics · 2025Article
- Advancing the Accuracy of Anti-MRSA Peptide Prediction Through Integrating Multi-Source Protein Language Models.Interdisciplinary sciences, computational life sciences · 2025Article
- MegSite: an accurate nucleic acid-binding residue prediction method based on multimodal protein language model.Briefings in bioinformatics · 2025Article
- Predicting nucleic acid binding sites by attention map-guided graph convolutional network with protein language embeddings and physicochemical information.Briefings in bioinformatics · 2025Article
- MESM: integrating multi-source data for high-accuracy protein-protein interactions prediction through multimodal language models.BMC biology · 2025Article
- Article
- Machine Learning for Protein Function Prediction.Methods in molecular biology (Clifton, N.J.) · 2025Review
- iProtDNA-SMOTE: Enhancing protein-DNA binding sites prediction through imbalanced graph neural networks.PloS one · 2025Article
- Advances in Language-Model-Informed Protein-Nucleic Acid Binding Site Prediction.Methods in molecular biology (Clifton, N.J.) · 2025Article
- Twenty years of advances in prediction of nucleic acid-binding residues in protein sequences.Briefings in bioinformatics · 2024Review
- HemoFuse: multi-feature fusion based on multi-head cross-attention for identification of hemolytic peptides.Scientific reports · 2024Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
Abstract
Efficient and accurate recognition of protein-DNA interactions is vital for understanding the molecular mechanisms of related biological processes and further guiding drug discovery. Although the current experimental protocols are the most precise way to determine protein-DNA binding sites, they tend to be labor-intensive and time-consuming. There is an immediate need to design efficient computational approaches for predicting DNA-binding sites. Here, we proposed ULDNA, a new deep-learning model, to deduce DNA-binding sites from protein sequences. This model leverages an LSTM-attention architecture, embedded with three unsupervised language models that are pre-trained on large-scale sequences from multiple database sources. To prove its effectiveness, ULDNA was tested on 229 protein chains with experimental annotation of DNA-binding sites. Results from computational experiments revealed that ULDNA significantly improves the accuracy of DNA-binding site prediction in comparison with 17 state-of-the-art methods. In-depth data analyses showed that the major strength of ULDNA stems from employing three transformer language models. Specifically, these language models capture complementary feature embeddings with evolution diversity, in which the complex DNA-binding patterns are buried. Meanwhile, the specially crafted LSTM-attention network effectively decodes evolution diversity-based embeddings as DNA-binding results at the residue level. Our findings demonstrated a new pipeline for predicting DNA-binding sites on a large scale with high accuracy from protein sequence alone.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.