Evidence map›Paper›PMID 42101927›Full record

ArticleBriefings in bioinformatics2026

When complexity does not pay: benchmarking deep learning and ensemble methods for biomarker discovery.

Cyrille Mesue Njume, Irene Petracci, Sonia Bellini, Katarzyna Goljanek-Whysall, Leo R Quinlan, Agnieszka Fiszer, Barbara Borroni, Roberta Ghidoni, Asli Kumbasar, Ali Cakmak

Abstract read
In one paragraph

Article in Briefings in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Cyrille Mesue NjumeDepartment of Molecular Biology and Genetics, Ayazaga Campus, Istanbul Technical University, Reşitpaşa, Sarıyer, 34467 Istanbul, Turkey.ORCID 0009-0001-6527-9108
Irene PetracciMolecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, 25125 Brescia, Italy.
Sonia BelliniMolecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, 25125 Brescia, Italy.ORCID 0000-0002-6162-252X
Katarzyna Goljanek-WhysallDiscipline of Physiology, School of Medicine, University of Galway, H91 TH33 Galway, Ireland.
Leo R QuinlanDiscipline of Physiology, School of Medicine, University of Galway, H91 TH33 Galway, Ireland.
Agnieszka FiszerDepartment of Medical Biotechnology, Institute of Bioorganic Chemistry, Polish Academy of Sciences, Noskowskiego 12/14, 61-704 Poznan, Poland.
Barbara BorroniMolecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, 25125 Brescia, Italy.
Roberta GhidoniMolecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, 25125 Brescia, Italy.
Asli KumbasarDepartment of Molecular Biology and Genetics, Ayazaga Campus, Istanbul Technical University, Reşitpaşa, Sarıyer, 34467 Istanbul, Turkey.
Ali CakmakDepartment of Computer Engineering, Ayazaga Campus, Istanbul Technical University, Reşitpaşa, Sarıyer, 34467 Istanbul, Turkey.ORCID 0000-0002-1382-6130

Funding

EU Joint Programme-Neurodegenerative Disease Research 124N069Health Research Board JPND-2023-1Italian Ministry of Health-EU Joint Programme-Neurodegenerative Disease ResearchNational Center for High-Performance Computing 1009742021Polish National Science Centre 2023/05/Y/NZ3/00160Scientific and Technological Research Council of TurkeyScientific Research Projects Unit of Istanbul Technical University TGA-2025-46998
6 · The paper itself

Abstract

The integration of multi-omics data holds great promise for identifying robust and clinically relevant biomarkers, yet the increasing complexity of computational methods raises questions about their practical utility. In this study, we present a comprehensive benchmarking framework that evaluates 27 feature selection strategies and 11 predictive models across three real-world disease cohorts: Alzheimer's disease, progressive supranuclear palsy, and breast cancer. We compare traditional machine learning, ensemble-based methods, and state-of-the-art deep learning models in terms of predictive performance, stability, and biological interpretability. Our results reveal that ensemble feature selection consistently improves robustness and accuracy, particularly for compact biomarker panels. Surprisingly, deep learning models did not outperform simpler classifiers such as logistic regression (L.Regression), support vector machines, or multilayer perceptrons, which often achieved comparable or superior results with lower computational cost and greater interpretability. Triple-omics yielded the highest validation, followed by dual-omics and then single-omics (Triple > Dual > Single). Biological validation against five independent databases confirmed the clinical relevance of the identified biomarkers, including both well-established and novel candidates. To support reproducibility and community adoption, we provide a web-based tool for applying our benchmarking pipeline. Our findings advocate for a pragmatic approach to biomarker discovery-prioritizing methodological transparency, reproducibility, and biological insight over algorithmic complexity.

Indexed as

BiomarkersComputational BiologyDeep LearningAlzheimer DiseaseBenchmarkingBreast NeoplasmsEnsemble LearningHumansMultiomicsReproducibility of ResultsBiomarkersbiomarker discoveryensemble rank aggregationfeature selection benchmarkingintegrative bioinformaticsmulti-omics integration

Identifiers

PMID42101927
PMCPMC13155123

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.