Evidence map›Paper›PMID 42487110›Full record

ArticleBMC bioinformatics2026

Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2.

Zohreh Piroozeh, Ildem Akerman, Olga V Kalinina, Stefan Kesselheim, Alina Bazarova

Abstract read
In one paragraph

Article in BMC bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Zohreh PiroozehJülich Supercomputing Centre, Forschungszentrum Jülich, Jülich, Germany. z.piroozeh@fz-juelich.de.
Ildem AkermanInstitute of Biomedical Research, University of Birmingham, Birmingham, UK.
Olga V KalininaCentre for Bioinformatics, Saarland University, Saarbrücken, Germany.
Stefan KesselheimJülich Supercomputing Centre, Forschungszentrum Jülich, Jülich, Germany.
Alina BazarovaJülich Supercomputing Centre, Forschungszentrum Jülich, Jülich, Germany. al.bazarova@fz-juelich.de.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundDNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals.

resultsWe evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2's learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition.

conclusionOur study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2's alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting.

Indexed as

Computational BiologyDNA ReplicationReplication OriginSaccharomyces cerevisiaeDNA, FungalDNA, FungalBudding yeastDNABERTDNA replication originsTokenizationXAI

Identifiers

PMID42487110
PMCPMC13393499

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.