Evidence map›Paper›PMID 39897946›Full record

ArticleBioinformatics advances2025

A comparative analysis of gene expression profiling by statistical and machine learning approaches.

Myriam Bontonou, Anaïs Haget, Maria Boulougouri, Benjamin Audit, Pierre Borgnat, Jean-Michel Arbona

Abstract read
In one paragraph

Article in Bioinformatics advances, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Article
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Myriam BontonouCNRS, ENS de Lyon, Inserm, LBMC, UMR5239, U1293, F-69342 Lyon Cedex 07, France.ORCID https://orcid.org/0000-0002-0010-5457
Anaïs HagetLTS2 Laboratory, EPFL, 1015 Lausanne, Switzerland.
Maria BoulougouriLTS2 Laboratory, EPFL, 1015 Lausanne, Switzerland.
Benjamin AuditCNRS, ENS de Lyon, LPENSL, UMR5672, F-69342 Lyon Cedex 07, France.ORCID https://orcid.org/0000-0003-2683-9990
Pierre BorgnatCNRS, ENS de Lyon, LPENSL, UMR5672, F-69342 Lyon Cedex 07, France.ORCID https://orcid.org/0000-0003-4536-8354
Jean-Michel ArbonaCNRS, ENS de Lyon, Inserm, LBMC, UMR5239, U1293, F-69342 Lyon Cedex 07, France.ORCID https://orcid.org/0000-0001-6166-9056

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Motivation: Many machine learning (ML) models developed to classify phenotype from gene expression data provide interpretations for their decisions, with the aim of understanding biological processes. For many models, including neural networks, interpretations are lists of genes ranked by their importance for the predictions, with top-ranked genes likely linked to the phenotype. In this article, we discuss the limitations of such approaches using integrated gradient, an explainability method developed for neural networks, as an example. Results: Experiments are performed on RNA sequencing data from public cancer databases. A collection of ML models, including multilayer perceptrons and graph neural networks, are trained to classify samples by cancer type. Gene rankings from integrated gradients are compared to genes highlighted by statistical feature selection methods such as DESeq2 and other learning methods measuring global feature contribution. Experiments show that a small set of top-ranked genes is sufficient to achieve good classification. However, similar performance is possible with lower-ranked genes, although larger sets are required. Moreover, significant differences in top-ranked genes, especially between statistical and learning methods, prevent a comprehensive biological understanding. In conclusion, while these methods identify pathology-specific biomarkers, the completeness of gene sets selected by explainability techniques for understanding biological processes remains uncertain. Availability and implementation: Python code and datasets are available at https://github.com/mbonto/XAI_in_genomics.

Identifiers

PMID39897946
PMCPMC11783302

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.