Evidence map›Paper›PMID 38364855›Full record

ArticleNucleic acids research2024

Prediction of DNA i-motifs via machine learning.

Bibo Yang, Dilek Guneri, Haopeng Yu, Elisé P Wright, Wenqian Chen, Zoë A E Waller, Yiliang Ding

Abstract read
In one paragraph

Article in Nucleic acids research, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.

0numbers the graph read from it
0cells of the map it votes in
15citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

15 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Review
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Review
  11. Article
  12. Article
  13. Review
  14. Article
  15. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Bibo YangDepartment of Cell and Developmental Biology, John Innes Centre, Norwich Research Park, Norwich NR4 7UH, UK.
Dilek GuneriSchool of Pharmacy, University College London, London WC1N 1AX, UK.
Haopeng YuDepartment of Cell and Developmental Biology, John Innes Centre, Norwich Research Park, Norwich NR4 7UH, UK.ORCID 0000-0002-5184-2430
Elisé P WrightMolecular Physiology School of Medicine, and Molecular Medicine Research Group, University of Western Sydney, Campbelltown, NSW 1797, Australia.
Wenqian ChenSchool of Pharmacy, University College London, London WC1N 1AX, UK.
Zoë A E WallerSchool of Pharmacy, University College London, London WC1N 1AX, UK.ORCID 0000-0001-8538-0484
Yiliang DingDepartment of Cell and Developmental Biology, John Innes Centre, Norwich Research Park, Norwich NR4 7UH, UK.ORCID 0000-0003-4161-6365

Funding

BBSRC BB/X01102X/1BBSRC Horizon Europe Guarantee EP/Y009886/1BBSRC Norwich Research Park Biosciences Doctoral Training Partnership 2578674Biotechnology and Biological Sciences Research Council BB/W000962/1Biotechnology and Biological Sciences Research Council BB/X01102X/1Human Frontier Science Program Fellowship LT001077/2021-L
6 · The paper itself

Abstract

i-Motifs (iMs), are secondary structures formed in cytosine-rich DNA sequences and are involved in multiple functions in the genome. Although putative iM forming sequences are widely distributed in the human genome, the folding status and strength of putative iMs vary dramatically. Much previous research on iM has focused on assessing the iM folding properties using biophysical experiments. However, there are no dedicated computational tools for predicting the folding status and strength of iM structures. Here, we introduce a machine learning pipeline, iM-Seeker, to predict both folding status and structural stability of DNA iMs. The programme iM-Seeker incorporates a Balanced Random Forest classifier trained on genome-wide iMab antibody-based CUT&Tag sequencing data to predict the folding status and an Extreme Gradient Boosting regressor to estimate the folding strength according to both literature biophysical data and our in-house biophysical experiments. iM-Seeker predicts DNA iM folding status with a classification accuracy of 81% and estimates the folding strength with coefficient of determination (R2) of 0.642 on the test set. Model interpretation confirms that the nucleotide composition of the C-rich sequence significantly affects iM stability, with a positive correlation with sequences containing cytosine and thymine and a negative correlation with guanine and adenine.

Indexed as

DNAMachine LearningNucleotide MotifsBase SequenceCytosineHumansCytosineDNA

Identifiers

PMID38364855
PMCPMC10954440

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.