Evidence map›Paper›PMID 40672173›Full record

ArticlebioRxiv : the preprint server for biology2025

mRNABench: A curated benchmark for mature mRNA property and function prediction.

Ruian Ian Shi, Taykhoom Dalal, Philip Fradkin, Divya Koyyalagunta, Simran Chhabria, Andrew Jung, Cyrus Tam, Defne Ceyhan, Jessica Lin, Kaitlin U Laverty and 3 more

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors.

Ruian Ian ShiDepartment of Computer Science, University of Toronto.
Taykhoom DalalComputational and Systems Biology Program, Sloan Kettering Institute.
Philip FradkinDepartment of Computer Science, University of Toronto.
Divya KoyyalaguntaComputational and Systems Biology Program, Sloan Kettering Institute.
Simran ChhabriaComputational and Systems Biology Program, Sloan Kettering Institute.
Andrew JungDepartment of Electrical and Computer Engineering, University of Toronto.
Cyrus TamComputational and Systems Biology Program, Sloan Kettering Institute.ORCID 0009-0006-5545-5557
Defne CeyhanComputational and Systems Biology Program, Sloan Kettering Institute.
Jessica LinComputational and Systems Biology Program, Sloan Kettering Institute.
Kaitlin U LavertyVector Institute.ORCID 0009-0004-4341-4665
Ilyes BaaliComputational and Systems Biology Program, Sloan Kettering Institute.
Bo WangDepartment of Computer Science, University of Toronto.ORCID 0000-0002-9620-3413
Quaid MorrisComputational and Systems Biology Program, Sloan Kettering Institute.ORCID 0000-0002-2760-6999

Funding

X-RAY CRYSTALLOGRAPHYP30CA008748 · NCI · SLOAN-KETTERING INSTITUTE FOR CANCER RES · PI SELWYN M VICKERS · 1985 to 2026
$347.4M
Post-transcriptional Regulatory NetworksR01HG013328 · NHGRI · SLOAN-KETTERING INST CAN RESEARCH · PI Quaid Morris · 2023 to 2026
$2.6M
NCI NIH HHS P30 CA008748NHGRI NIH HHS R01 HG013328
6 · The paper itself

Abstract

Messenger RNA (mRNA) is central in gene expression, and its half-life, localization, and translation efficiency drive phenotypic diversity in eukaryotic cells. While supervised learning has widely been used to study the mRNA regulatory code, self-supervised foundation models support a wider range of transfer learning tasks. However, the dearth and homogeneity of standardized benchmarks limit efforts to pinpoint the strengths of various models. Here, we present mRNABench, a comprehensive benchmarking suite for mature mRNA biology that evaluates the representational quality of mature mRNA embeddings from self-supervised nucleotide foundation models. We curate ten datasets and 59 prediction tasks that broadly capture salient properties of mature mRNA, and assess the performance of 18 families of nucleotide foundation models for a total of 135K experiments. Using these experiments, we study parameter scaling, compositional generalization from learned biological features, and correlations between sequence compressibility and performance. We identify synergies between two self-supervised learning objectives, and pre-train a new Mamba-based model that achieves state-of-the-art performance using 700x fewer parameters. mRNABench can be found at: https://github.com/morrislab/mRNABench.

Identifiers

PMID40672173
PMCPMC12265608

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.