Evidence map›Paper›PMID 39975320›Full record

ArticlebioRxiv : the preprint server for biology2025

Principled PCA separates signal from noise in omics count data.

Jay S Stanley, Junchen Yang, Ruiqi Li, Ofir Lindenbaum, Dmitry Kobak, Boris Landa, Yuval Kluger

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Jay S StanleyProgram in Applied Mathematics, Yale University, New Haven, CT, USA.ORCID 0000-0002-9227-8928
Junchen YangInterdepartmental Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA.ORCID 0000-0003-0988-1564
Ruiqi LiInterdepartmental Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA.
Ofir LindenbaumFaculty of Engineering, Bar Ilan University, Ramat-Gan, Israel.
Dmitry KobakHertie Institute for AI in Brain Health, University of Tübingen, Germany.ORCID 0000-0002-5639-7209
Boris LandaProgram in Applied Mathematics, Yale University, New Haven, CT, USA.
Yuval KlugerProgram in Applied Mathematics, Yale University, New Haven, CT, USA.ORCID 0000-0002-3035-071X

Funding

Yale SPORE in Skin CancerP50CA121974 · NCI · YALE UNIVERSITY · PI MARCUS W BOSENBERG, Harriet M. Kluger · 2006 to 2026
$43.9M
The Y-SCORCH Data Generation Center at Yale for Single-Cell Opioid Responses in the Context of HIVUM1DA051410 · NIDA · YALE UNIVERSITY · PI GERSTEIN, MARK BENDER, KLUGER, YUVAL · 2020 to 2025
$15.3M
M-SCORCH: Methamphetamine use disorder data generation center for Single Cell Opioid Responses in the Context of HIVU01DA053628 · NIDA · YALE UNIVERSITY · PI HO, YA-CHI, SESTAN, NENAD · 2021 to 2025
$9.5M
Yale TMC for Cellular Senescence in Lymphoid OrgansU54AG076043 · NIA · YALE UNIVERSITY · PI FAN, RONG, HALENE, STEPHANIE · 2021 to 2025
$7.0M
Yale Murine-TMC on Immune Cell Senescence Derived InflammationU54AG079759 · NIA · YALE UNIVERSITY · PI DIXIT, VISHWA DEEP, MONTGOMERY, RUTH R · 2022 to 2025
$6.5M
Evaluating the role of opioid medication assisted therapies in HIV-1 Persistence for persons living with HIV and opioid use disordersR33DA047037 · NIDA · YALE UNIVERSITY · PI HO, YA-CHI, KLUGER, YUVAL · 2021 to 2022
$1.7M
EFFICIENT METHODS FOR CALIBRATION, CLUSTERING, VISUALIZATION AND IMPUTATION OF LARGE scRNA-seq DATAR01GM131642 · NIGMS · YALE UNIVERSITY · PI KLUGER, YUVAL · 2019 to 2022
$1.6M
NCI NIH HHS P50 CA121974NIA NIH HHS U54 AG076043NIA NIH HHS U54 AG079759NIDA NIH HHS R33 DA047037NIDA NIH HHS U01 DA053628NIDA NIH HHS UM1 DA051410NIGMS NIH HHS R01 GM131642
6 · The paper itself

Abstract

Principal component analysis (PCA) is indispensable for processing high-throughput omics datasets, as it can extract meaningful biological variability while minimizing the influence of noise. However, the suitability of PCA is contingent on appropriate normalization and transformation of count data, and accurate selection of the number of principal components; improper choices can result in the loss of biological information or corruption of the signal due to excessive noise. Typical approaches to these challenges rely on heuristics that lack theoretical foundations. In this work, we present Biwhitened PCA (BiPCA), a theoretically grounded framework for rank estimation and data denoising across a wide range of omics modalities. BiPCA overcomes a fundamental difficulty with handling count noise in omics data by adaptively rescaling the rows and columns - a rigorous procedure that standardizes the noise variances across both dimensions. Through simulations and analysis of over 100 datasets spanning seven omics modalities, we demonstrate that BiPCA reliably recovers the data rank and enhances the biological interpretability of count data. In particular, BiPCA enhances marker gene expression, preserves cell neighborhoods, and mitigates batch effects. Our results establish BiPCA as a robust and versatile framework for high-throughput count data analysis.

Identifiers

PMID39975320
PMCPMC11838471

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.