Evidence map›Paper›PMID 38349057›Full record

ArticleBriefings in bioinformatics2024

ULDNA: integrating unsupervised multi-source language models with LSTM-attention network for high-accuracy protein-DNA binding site prediction.

Yi-Heng Zhu, Zi Liu, Yan Liu, Zhiwei Ji, Dong-Jun Yu

Abstract read
In one paragraph

Article in Briefings in bioinformatics, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 21 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
21citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

21 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Machine Learning for Protein Function Prediction.Methods in molecular biology (Clifton, N.J.) · 2025
    Review
  17. Article
  18. Article
  19. Review
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Yi-Heng ZhuCollege of Artificial Intelligence, Nanjing Agricultural University, Nanjing 210095, China.
Zi LiuSchool of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China.
Yan LiuSchool of Information Engineering, Yangzhou University, Yangzhou 225000, China.ORCID 0000-0002-5331-3655
Zhiwei JiCollege of Artificial Intelligence, Nanjing Agricultural University, Nanjing 210095, China.ORCID 0000-0002-0891-7118
Dong-Jun YuSchool of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China.ORCID 0000-0002-6786-8053

Funding

Foundation of National Defense Key Laboratory of Science and Technology JZX7Y202001SY000901Jiangsu Funding Program for Excellent Postdoctoral Talent 2023ZB224National Natural Science Foundation of China 62372234Natural Science Foundation of Jiangsu BK20201304
6 · The paper itself

Abstract

Efficient and accurate recognition of protein-DNA interactions is vital for understanding the molecular mechanisms of related biological processes and further guiding drug discovery. Although the current experimental protocols are the most precise way to determine protein-DNA binding sites, they tend to be labor-intensive and time-consuming. There is an immediate need to design efficient computational approaches for predicting DNA-binding sites. Here, we proposed ULDNA, a new deep-learning model, to deduce DNA-binding sites from protein sequences. This model leverages an LSTM-attention architecture, embedded with three unsupervised language models that are pre-trained on large-scale sequences from multiple database sources. To prove its effectiveness, ULDNA was tested on 229 protein chains with experimental annotation of DNA-binding sites. Results from computational experiments revealed that ULDNA significantly improves the accuracy of DNA-binding site prediction in comparison with 17 state-of-the-art methods. In-depth data analyses showed that the major strength of ULDNA stems from employing three transformer language models. Specifically, these language models capture complementary feature embeddings with evolution diversity, in which the complex DNA-binding patterns are buried. Meanwhile, the specially crafted LSTM-attention network effectively decodes evolution diversity-based embeddings as DNA-binding results at the residue level. Our findings demonstrated a new pipeline for predicting DNA-binding sites on a large scale with high accuracy from protein sequence alone.

Indexed as

Data AnalysisLanguageAmino Acid SequenceBinding SitesDatabases, Factualdeep learningevolution diversityLSTM-attention networkprotein–DNA interactionunsupervised protein language model

Identifiers

PMID38349057
PMCPMC10939370

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.