Evidence map›Paper›PMID 41920894›Full record

ArticlePLoS computational biology2026

Predictive modeling of gene expression and localization of DNA binding site using deep convolutional neural networks.

Arman Karshenas, Tom Röschinger, Hernan G Garcia

Erratum issuedAbstract read
In one paragraph

Article in PLoS computational biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

3 authors.

Arman KarshenasBiophysics Graduate Group, University of California at Berkeley, Berkeley, California, United States of America.ORCID https://orcid.org/0000-0001-5477-1861
Tom RöschingerDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, California, United States of America.ORCID https://orcid.org/0000-0002-4900-3216
Hernan G GarciaBiophysics Graduate Group, University of California at Berkeley, Berkeley, California, United States of America.ORCID https://orcid.org/0000-0002-5212-3649

Funding

Predictive understanding of the temporal control of transcription in Drosophila developmentR01GM139913 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI GARCIA, HERNAN GUSTAVO · 2021 to 2024
$1.2M
Physical Biology of Transcriptional Control in Embryonic DevelopmentR35GM158200 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI Hernan Gustavo Garcia · 2025 to 2026
$847k
In vivo mechanisms of Dorsal-mediated transcriptional control in developmentR01GM152815 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI GARCIA, HERNAN GUSTAVO · 2024 to 2024
$306k
NIGMS NIH HHS R01 GM139913NIGMS NIH HHS R01 GM152815NIGMS NIH HHS R35 GM158200
6 · The paper itself

Abstract

Despite the sequencing revolution, large swaths of the genomes sequenced to date lack any information about the arrangement of transcription factor binding sites on regulatory DNA. Massively Parallel Reporter Assays (MPRAs) have the potential to dramatically accelerate our genomic annotations by making it possible to measure the gene expression levels driven by thousands of mutational variants of a regulatory region. However, the interpretation of such data often assumes that each base pair in a regulatory sequence contributes independently to the overall gene expression. To enable the analysis of this data in a manner that accounts for possible correlations between distant bases along a regulatory sequence, we developed the Deep learning Adaptable Regulatory Sequence Identifier (DARSI). This convolutional neural network leverages MPRA data for training specific models for each operon to predict gene expression levels directly from raw regulatory DNA sequences. By harnessing this predictive capacity, DARSI systematically identifies transcription factor binding sites within regulatory regions at single-base pair resolution. To validate its predictions, we benchmarked DARSI against curated databases, confirming its accuracy in predicting known transcription factor binding sites. Additionally, DARSI predicted novel unmapped binding sites, paving the way for future experimental efforts to confirm the existence of these binding sites and to identify the transcription factors that target those sites. Thus, DARSI provides a new framework for MPRA experimental data analysis, it generates experimentally actionable predictions that can feed iterations of the theory-experiment cycle aimed at reaching a predictive understanding of transcriptional control. Here, we developed a deep learning approach-called DARSI-that leverages these massively parallel reporter assays to predict levels of gene expression from DNA sequences and help locate these important binding sites. By training our model to recognize DNA sequence patterns that affect gene expression, our method not only finds known binding sites with high accuracy, but also predicts new binding sites that call for future experimental scrutiny.

Indexed as

DNAGene ExpressionBinding SitesComputational BiologyConvolutional Neural NetworksDeep LearningModels, GeneticPredictive Learning ModelsRegulatory Sequences, Nucleic AcidTranscription FactorsDNATranscription Factors

Identifiers

PMID41920894
PMCPMC13052891

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.