Evidence map›Paper›PMID 40702706›Full record

ArticleBriefings in bioinformatics2025

Benchmarking transcription factor binding site prediction models: a comparative analysis on synthetic and biological data.

Manuel Tognon, Alisa Kumbara, Andrea Betti, Lorenzo Ruggeri, Rosalba Giugno

Abstract readComparative Study
In one paragraph

Article in Briefings in bioinformatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Review
  2. Toward mechanistic virtual immune cells.Nature biotechnology · 2026
    Article
  3. Review
  4. Article
  5. Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Manuel TognonComputer Science Department, University of Verona, Strada Le Grazie 15, Verona, VR 37134, Italy.ORCID 0000-0002-6707-2071
Alisa KumbaraDepartment of Engineering for Innovation Medicine, University of Verona, Strada Le Grazie 15, Verona, VR 37134, Italy.ORCID 0009-0002-3899-3928
Andrea BettiComputer Science Department, University of Verona, Strada Le Grazie 15, Verona, VR 37134, Italy.ORCID 0009-0006-8822-1677
Lorenzo RuggeriComputer Science Department, University of Verona, Strada Le Grazie 15, Verona, VR 37134, Italy.ORCID 0009-0000-2516-1175
Rosalba GiugnoComputer Science Department, University of Verona, Strada Le Grazie 15, Verona, VR 37134, Italy.ORCID 0000-0001-9843-7638

Funding

European Union - NextGenerationEU
6 · The paper itself

Abstract

Transcription factors (TFs) are essential regulatory proteins controlling the cellular transcriptional states by binding to specific DNA sequences known as transcription factor binding sites (TFBSs) or motifs. Accurate TFBS identification is crucial for unraveling regulatory mechanisms driving cellular dynamics. Over the years, various computational approaches have been developed to model TFBSs, with position weight matrices (PWMs) being one of the most widely adopted methods. PWMs provide a probabilistic framework by representing nucleotide frequencies at every position within the binding site. While effective and interpretable, PWMs face significant limitations, such as their inability to capture positional dependencies or model complex interactions. To address these, advanced methods, like support vector machine (SVM)-based, and deep learning (DL)-based models, have been introduced. Leveraging human ChIP-seq data from ENCODE, we systematically benchmarked the predictive performance of PWM, SVM-, and DL-based models across different scenarios. We evaluate the impact of key factors such as training dataset size, sequence length, and kernel functions (for SVMs) on models' performance. Additionally, we explore the impact of synthetic versus real biological background data during model training. Our analysis highlights strengths and limitations of each approach under different conditions, providing practical guidance for selecting and tailoring models to specific biological datasets. To complement our analysis, we present a comprehensive database of pretrained SVM models for TFBS detection, trained on human ChIP-seq data from diverse cell lines and tissues. This resource aims to facilitate broader adoption of SVM-based methods in TFBS prediction and enhance their practical utility in regulatory genomics research.

Indexed as

Computational BiologyTranscription FactorsBenchmarkingBinding SitesChromatin Immunoprecipitation SequencingDeep LearningHumansSupport Vector MachineTranscription Factorsbioinformaticsepigeneticsmachine learningtranscription factor binding

Identifiers

PMID40702706
PMCPMC12286778

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.