Evidence map›Paper›PMID 41279114›Full record

ArticlebioRxiv : the preprint server for biology2026

Scaling Large Language Models for Next-Generation Single-Cell Analysis.

Syed Asad Rizvi, Daniel Levine, Aakash Patel, Shiyang Zhang, Eric Wang, Curtis Jamison Perry, Ivan Vrkic, Nicole Mayerli Constante, Zirui Fu, Sizhuang He and 18 more

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

28 authors.

Syed Asad RizviYale University, Google Research.ORCID 0000-0002-7932-9524
Daniel LevineYale University.
Aakash PatelYale University.
Shiyang ZhangYale University.
Eric WangGoogle DeepMind.
Curtis Jamison PerryYale University.
Ivan VrkicYale University.
Nicole Mayerli ConstanteYale University.
Zirui FuYale University.
Sizhuang HeYale University.
David ZhangYale University.
Cerise TangYale University.
Zhuoyang LyuBrown University.
Rayyan DarjiYale University.
Chang LiYale University.
Emily SunYale University.
David JeongYale University.
Lawrence ZhaoYale University.
Jennifer KwanYale University.
David BraunYale University.
Brian HaflerYale University.
Hattie ChungYale University.ORCID 0000-0001-8417-5606
Rahul M DhodapkarUniversity of Southern California.ORCID 0000-0002-2014-7515
Paul JaegerGoogle DeepMind.
Bryan PerozziGoogle Research.
Jeffrey IshizukaYale University.ORCID 0000-0002-4271-7312
Shekoofeh AziziGoogle DeepMind.ORCID 0000-0002-7447-6031
David van DijkYale University.ORCID 0000-0003-3911-9925

Funding

An integrative, data-driven, and computational approach to uncovering dynamic mechanisms of early viral infectionR35GM143072 · NIGMS · YALE UNIVERSITY · PI VAN DIJK, DAVID · 2021 to 2025
$2.1M
Targeting the inflammatory response in age-related macular degenerationR01EY034234 · NEI · YALE UNIVERSITY · PI Brian P Hafler · 2022 to 2026
$2.0M
NEI NIH HHS R01 EY034234NIGMS NIH HHS R35 GM143072
6 · The paper itself

Abstract

Single-cell RNA sequencing has transformed our understanding of cellular diversity, yet current single-cell foundation models (scFMs) remain limited in their scalability, flexibility across diverse tasks, and ability to natively integrate textual information. In this work, we build upon the Cell2Sentence (C2S) framework, which represents scRNA-seq profiles as textual "cell sentences," to train Large Language Models (LLMs) on a corpus comprising over one billion tokens of transcriptomic data, biological text, and metadata. Scaling the model to 27 billion parameters yields consistent improvements in predictive and generative capabilities and supports advanced downstream tasks that require synthesis of information across multi-cellular contexts. Targeted fine-tuning with modern reinforcement learning techniques produces strong performance in perturbation response prediction, natural language interpretation, and complex biological reasoning. This predictive strength enabled a dual-context virtual screen that nominated the kinase inhibitor silmitasertib (CX-4945) as a candidate for context-selective upregulation of antigen presentation. Experimental assessment in human cell models unseen during training supported this prediction, demonstrating that C2S-Scale can effectively guide the discovery of context-conditioned biology. C2S-Scale unifies transcriptomic and textual data at unprecedented scales, surpassing both specialized single-cell models and general-purpose LLMs to provide a platform for next-generation single-cell analysis and the development of "virtual cells."

Identifiers

PMID41279114
PMCPMC12632461

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.