Evidence map›Paper›PMID 40501927›Full record

ArticlebioRxiv : the preprint server for biology2025

"Frustratingly easy" domain adaptation for cross-species transcription factor binding prediction.

Mark Maher Ebeid, Ali Tuğrul Balcı, Maria Chikina, Panayiotis V Benos, Dennis Kostka

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Mark Maher EbeidDepartment of Computational & Systems Biology University of Pittsburgh School of Medicine and Joint Carnegie Mellon-University of Pittsburgh Ph.D. Program in Computational Biology, University of Pittsburgh, Pittsburgh, PA, USA.
Ali Tuğrul BalcıDepartment of Computational & Systems Biology University of Pittsburgh School of Medicine and Joint Carnegie Mellon-University of Pittsburgh Ph.D. Program in Computational Biology, University of Pittsburgh, Pittsburgh, PA, USA.
Maria ChikinaDepartment of Computational & Systems Biology University of Pittsburgh School of Medicine and Joint Carnegie Mellon-University of Pittsburgh Ph.D. Program in Computational Biology, University of Pittsburgh, Pittsburgh, PA, USA.
Panayiotis V BenosDepartment of Computational & Systems Biology University of Pittsburgh School of Medicine and Joint Carnegie Mellon-University of Pittsburgh Ph.D. Program in Computational Biology, University of Pittsburgh, Pittsburgh, PA, USA.ORCID 0000-0003-3172-3132
Dennis KostkaDepartment of Computational & Systems Biology University of Pittsburgh School of Medicine and Joint Carnegie Mellon-University of Pittsburgh Ph.D. Program in Computational Biology, University of Pittsburgh, Pittsburgh, PA, USA.ORCID 0000-0002-1460-5487

Funding

Genomic Analysis of Tissue and Cellular Heterogeneity in IPFR01HL127349 · NHLBI · YALE UNIVERSITY · PI BENOS, PANAGIOTIS V, KAMINSKI, NAFTALI · 2015 to 2025
$5.9M
PKR sensing of mitochondrial dsRNA in childhood Sjogrens diseaseR01DE032707 · NIDCR · UNIVERSITY OF FLORIDA · PI SEUNGHEE CHA · 2023 to 2026
$2.4M
Interpretable graphical models for large multi-modal COPD data (R01HL159805)R01HL159805 · NHLBI · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI BENOS, PANAGIOTIS V, SPIRTES, PETER · 2021 to 2024
$2.0M
NHLBI NIH HHS R01 HL127349NHLBI NIH HHS R01 HL159805NIDCR NIH HHS R01 DE032707
6 · The paper itself

Abstract

Motivation: Understanding how DNA sequence encodes gene regulation remains a central challenge in genomics. While deep learning models can predict regulatory activity from sequence with high accuracy, their generalizability across species-and thus their ability to capture fundamental biological principles-remains limited. Cross-species prediction provides a powerful test of model robustness and offers a window into conserved regulatory logic, but effectively bridging species-specific genomic differences remains a major barrier. Results: We present MORALE, a novel and scalable domain adaptation framework that significantly advances cross-species prediction of transcription factor (TF) binding. By aligning statistical moments of sequence embeddings across species, MORALE enables deep learning models to learn species-invariant regulatory features without requiring adversarial training or complex architectures. Applied to multi-species TF ChIP-seq datasets, MORALE achieves state-of-the-art performance-outperforming both baseline and adversarial approaches across all TFs-while preserving model interpretability and recovering canonical motifs with greater precision. In the five-species transfer setting, MORALE not only improves human prediction accuracy beyond human-only training but also reveals regulatory features conserved across mammals. These results highlight the potential of simple yet powerful domain adaptation techniques to drive generalization and discovery in regulatory genomics. Crucially, MORALE is architecture-agnostic and can be seamlessly integrated into any embedding-based sequence model. Availability: Code is available at https://github.com/loudrxiv/frustrating.

Identifiers

PMID40501927
PMCPMC12154900

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.