ArticleNature communications2024
Improving prediction performance of general protein language model by domain-adaptive pretraining on DNA-binding protein.
Article in Nature communications, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 26 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
26 citing papers in PubMed.
- Protein-DNA Binding Sites Prediction via Integrating Pretrained Large Language Models and Contrastive Learning.Interdisciplinary sciences, computational life sciences · 2026Article
- Machine Learning of Personal Repertoires From Public T Cell Receptors.Immunological reviews · 2026Review
- Integrating protein and DNA embeddings for improving genome-wide transcription factor binding site prediction.NAR genomics and bioinformatics · 2026Article
- Predicting protein-nucleic acid interactions via protein language models with biophysical and evolutionary priors.iScience · 2026Article
- QSyncFold: quantum neural network for multidimensional sync-discovery in protein folding.Briefings in bioinformatics · 2026Article
- Protein foundation models: a comprehensive survey.Science China. Life sciences · 2026Review
- A survey on large language models in biology and chemistry.Experimental & molecular medicine · 2026Review
- ESM-PsyPred: Leveraging Protein Language Models for Accurate Prediction of Psychrophilic Proteins.Interdisciplinary sciences, computational life sciences · 2026Article
- BiGKbhb: a bi-directional gated recurrent unit model for predicting lysine β-hydroxybutyrylation sites.BMC genomics · 2026Article
- Learning physical interactions to compose biological large language models.Communications chemistry · 2026Review
- Interpretable prediction of nucleic acid-binding proteins using a protein language model.Bioinformatics advances · 2026Article
- DBP-CanPred: a machine learning model for predicting cancer-causing mutations in DNA-binding proteins.Frontiers in bioinformatics · 2026Article
- Active learning-guided optimization of cell-free biosensors for lead testing in drinking water.Nature communications · 2025Article
- Unveiling the landscape of prokaryotic global regulators through deep protein language models.mSystems · 2025Article
- PPAC: Predicting Protein-Protein Affinity Changes Induced by Amino Acid Mutations Using Protein Large Language Models.ACS omega · 2025Article
- iDNA-DAPHA: a generic framework for methylation prediction via domain-adaptive pretraining and hierarchical attention.Briefings in bioinformatics · 2025Article
- Protein language model identifies disordered, conserved motifs implicated in phase separation.eLife · 2025Article
- LassoESM a tailored language model for enhanced lasso peptide property prediction.Nature communications · 2025Article
- MegSite: an accurate nucleic acid-binding residue prediction method based on multimodal protein language model.Briefings in bioinformatics · 2025Article
- Advancing the accuracy of clathrin protein prediction through multi-source protein language models.Scientific reports · 2025Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
DNA-protein interactions exert the fundamental structure of many pivotal biological processes, such as DNA replication, transcription, and gene regulation. However, accurate and efficient computational methods for identifying these interactions are still lacking. In this study, we propose a method ESM-DBP through refining the DNA-binding protein sequence repertory and domain-adaptive pretraining based the general protein language model. Our method considers the lacking exploration of general language model for DNA-binding protein domain-specific knowledge, so we screen out 170,264 DNA-binding protein sequences to construct the domain-adaptive language model. Experimental results on four downstream tasks show that ESM-DBP provides a better feature representation of DNA-binding protein compared to the original language model, resulting in improved prediction performance and outperforming the state-of-the-art methods. Moreover, ESM-DBP can still perform well even for those sequences with only a few homologous sequences. ChIP-seq on two predicted cases further support the validity of the proposed method.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.