Evidence map›Paper›PMID 40314439›Full record

ArticlemSystems2025

Compositional transformations can reasonably introduce phenotype-associated values into sparse features.

George I Austin, Tal Korem

Abstract read
In one paragraph

Article in mSystems, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Harnessing the microbiome for cancer therapy.Nature reviews. Microbiology · 2026
    Review
  2. Article
  3. Article
  4. Article
  5. Article
  6. Review
  7. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

2 authors.

George I AustinDepartment of Biomedical Informatics, Columbia University Irving Medical, New York, New York, USA.ORCID 0000-0002-7834-4968
Tal KoremProgram for Mathematical Genomics, Department of Systems Biology, Columbia University Irving Medical Center, New York, New York, USA.ORCID 0000-0002-0609-0858

Funding

Training in Biomedical Informatics at Columbia UniversityT15LM007079 · NLM · COLUMBIA UNIV NEW YORK MORNINGSIDE · PI NOEMIE ELHADAD, GEORGE M HRIPCSAK · 1992 to 2026
$28.9M
NLM NIH HHS T15 LM007079U.S. National Library of Medicine T15LM007079
6 · The paper itself

Abstract

Gihawi et al. (mBio 14:e01607-23, 2023, https://doi.org/10.1128/mbio.01607-23) argued that the analysis of tumor-associated microbiome data by Poore et al. (Nature 579:567-574, 2020, https://doi.org/10.1038/s41586-020-2095-1) is invalid because features that were originally very sparse (genera with mostly zero read counts) became associated with the phenotype following batch correction. Here, we examine whether such an observation should necessarily indicate issues with processing or machine learning pipelines. We show counterexamples using the centered log ratio (CLR) transformation, which is often used for analysis of compositional microbiome data. The CLR transformation has similarities to voom-SNM, the batch-correction method brought into question by Gihawi et al., and yet is a sample-wise operation that cannot, in itself, "leak" information or invalidate downstream analyses. We show that because the CLR transformation divides each value by the geometric mean of its sample, common imputation strategies for missing or zero values result in transformed features that are associated with the geometric mean. Through analyses of both synthetic and vaginal microbiome data sets, we demonstrate that when the geometric mean is associated with a phenotype, sparse and CLR-transformed features will also become associated with it. We re-analyze features highlighted by Gihawi et al. and demonstrate that the phenomenon of sparse features becoming phenotype-associated can also be observed after a CLR transformation, which serves as a counterexample to the claim that such an observation necessarily means information leakage. While we do not intend to address other concerns regarding tumor microbiome analyses, validate Poore et al.'s results, or evaluate batch-correction pipelines, we conclude that because phenotype-associated features that were initially sparse can be created by a sample-wise transformation that cannot artifactually inflate machine learning performance, their detection is not independently sufficient to demonstrate information leakage in machine learning pipelines. Microbiome data are multivariate, and as such, a value of 0 carries a different meaning for each sample. Many transformations, including CLR and other batch-correction methods, are likewise multivariate, and, as these issues demonstrate, each individual feature should be interpreted with caution. IMPORTANCE: Gihawi et al. claim that finding that a transformation turned highly sparse (mostly zero) features into features that are associated with a phenotype is sufficient to conclude that there is information leakage and to invalidate an analysis. This claim has critical implications for both the debate regarding The Cancer Genome Atlas (TCGA) cancer microbiome analysis and for interpretation and evaluation of analyses in the microbiome field at large. We show by counterexamples and by reanalysis that such transformations can be valid.

Indexed as

MicrobiotaFemaleHumansMachine LearningPhenotypecompositional data analysisimputationmachine learningmicrobiome

Identifiers

PMID40314439
PMCPMC12090810

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.