Evidence map›Paper›PMID 42370288›Full record

ArticleResearch square2026

Heavy-chain immune repertoire sequencing enables language-model prediction of antigen-specific antibodies.

Karen Paco, Mariana Mendivil, Zihao Zhang, Sanaz Zebardast, Christian Davila, Ryan Mooney, Peace Olatoyinbo, Tristan Yang, Sebastian Bassi, Virginia Gonzales and 14 more

Abstract readPreprint
In one paragraph

Article in Research square, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

24 authors.

Karen PacoKeck Graduate Institute.
Mariana MendivilKeck Graduate Institute.
Zihao ZhangKeck Graduate Institute.
Sanaz ZebardastKeck Graduate Institute.
Christian DavilaKeck Graduate Institute.
Ryan MooneyPomona College.
Peace OlatoyinboKeck Graduate Institute.
Tristan YangKeck Graduate Institute.
Sebastian BassiToyoko Labs.
Virginia GonzalesToyoko Labs.
Eva ChenPomona College.
Faisal Bin AshrafUniversity of California, Riverside.
Isabel CondoriKeck Graduate Institute.
Jonathan FelixKeck Graduate Institute.
Rashid AlamKeck Graduate Institute.
Jordan LayKeck Graduate Institute.
Malkiat JohalPomona College.
Karine Le RochUniversity of California, Riverside.
Ilya TolstorukovKeck Graduate Institute.
Jeniffer HernandezKeck Graduate Institute.
Fernando Barroso Da SilvaUniversidade de São Paulo.
Stefano LonardiUniversity of California, Riverside.
Matthew SazinskyPomona College.
Animesh RayKeck Graduate Institute.

Funding

Rapid response for pandemics: single cell sequencing and deep learning to predict antibody sequences against an emerging antigenR01AI169543 · NIAID · KECK GRADUATE INST OF APPLIED LIFE SCIS · PI HERNANDEZ, JENIFFER BERTHA, LONARDI, STEFANO · 2021 to 2023
$3.1M
NIAID NIH HHS R01 AI169543
6 · The paper itself

Abstract

Rapidly identifying antigen-specific antibodies within complex B cell repertoires is important for therapeutic antibody discovery, vaccine development, disease surveillance, and immune condition monitoring, especially for emerging pandemics. RNA deep sequencing can rapidly provide antibody sequence candidates, but predicting their binding-specificity has remained difficult. Here we show that antigen-specific antibodies can be predicted directly from mRNA-derived heavy-chain V(D)J deep sequencing of unselected immune repertoires by parameter-efficient fine-tuning of the protein language model ESM-2. We have achieved high antigen recognition accuracies across antibodies specific for SARS-CoV-2 spike protein, influenza hemagglutinin, and HIV glycoprotein gp120 antigens, respectively. Our fine-tuned language model, Antigen Specificity Predictor, when applied to unsorted peripheral blood repertoires from immunized mice by single-cell deep sequencing, could predict specific B cell receptors at high frequency, which were then experimentally validated. A significant overlap was obtained with the predicted receptors when benchmarked on previously unseen human B cell receptor sequences identified by barcoding-enabled affinity selection. In bulk mRNA-sequences of human immune repertoires, the predicted antigen-specific B cells exhibited characteristics reminiscent of biology-aware learning. Our model's performance cannot be explained by sequence memorization. We establish that unselected heavy chain antibody sequences alone carry sufficient signal for repertoire-scale computational antibody discovery and immune profiling, thus motivating potential extension to autoantibody identification in cancer and autoimmune diseases.

Identifiers

PMID42370288
PMCPMC13308380

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.