Evidence map›Paper›PMID 42182279›Full record

ArticlebioRxiv : the preprint server for biology2026

IntegrateRigor: annotation-free integration optimization for cell identity recovery reveals cancer-immune interface niches.

Zhiqian Zhai, Changhu Wang, Chengfeng Jiang, Ziqi Rong, Jingyi Jessica Li

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Zhiqian ZhaiDepartment of Statistics and Data Science, University of California, Los Angeles, CA 90095.ORCID 0009-0001-3104-7472
Changhu WangBiostatistics Program, Public Health Science Division, Fred Hutchinson Cancer, Seattle, WA 98109.ORCID 0009-0002-8567-0961
Chengfeng JiangDepartment of Statistics and Data Science, University of California, Los Angeles, CA 90095.ORCID 0009-0009-5021-0747
Ziqi RongPaul G. Allen School of Computer Science and Engineering, University of Washington, Seattle, WA 98195.ORCID 0000-0003-3760-8450
Jingyi Jessica LiDepartment of Statistics and Data Science, University of California, Los Angeles, CA 90095.ORCID 0000-0002-9288-5648

Funding

Experimental-data-based in-silico data generation platform to improve the accuracy and reliability of single-cell and spatial omics data analysisR01HG014687 · NHGRI · FRED HUTCHINSON CANCER CENTER · PI LI, JINGYI JESSICA · 2025 to 2025
$2.4M
Statistical Methods for Elucidating Regulatory Mechanisms and Functional Impacts of Transcriptome Variation at Population and Single-Cell ScalesR35GM140888 · NIGMS · UNIVERSITY OF CALIFORNIA LOS ANGELES · PI LI, JINGYI JESSICA · 2021 to 2025
$2.0M
NHGRI NIH HHS R01 HG014687NIGMS NIH HHS R35 GM140888
6 · The paper itself

Abstract

Integrating single-cell and spatial transcriptomics data across batches is essential for recovering comparable cell identities-including cell types, subtypes, and states-as a prerequisite for downstream analyses in multi-condition and large-scale studies. This task remains challenging because between-batch variation removal often conflicts with cell identity preservation, and current methods typically rely on generic highly variable gene selection and lack principled metrics for hyperparameter tuning when cell identity annotations are unavailable. Together, these limitations often lead to over-integration, which merges biologically distinct cell identities, or under-integration, which leaves cells separated by batch rather than identity. Here we introduce IntegrateRigor, a data-driven, annotation-free, method-agnostic framework that optimizes integration specifically for reliable cell identity recovery across batches. IntegrateRigor first selects genes whose expression patterns are stable across batches using a gene-wise likelihood-based batch stability score, excluding batch-sensitive genes that can bias cell identity alignment during integration. It then identifies the optimal integration configuration across methods and hyperparameters by defining a dataset-level integration score that explicitly balances between-batch variation removal against cell identity preservation, without requiring prior annotations. In a colorectal cancer single-cell and spatial transcriptomics dataset, IntegrateRigor revealed previously uncharacterized cancer-immune interface niches in the tumor microenvironment that were masked by under-integration under default settings and by over-integration in previous literature. Across diverse datasets spanning multiple sources of between-batch variation, IntegrateRigor consistently improved cell identity recovery by mitigating both over-integration and under-integration across five state-of-the-art integration methods. By transforming integration from a heuristic preprocessing step into a statistically principled, dataset-adaptive procedure for cell identity recovery, IntegrateRigor improves the reproducibility and biological discovery power of large-scale single-cell and spatial transcriptomics analyses.

Identifiers

PMID42182279
PMCPMC13192717

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.