Evidence map›Paper›PMID 42680564›Full record

ArticleRNA (New York, N.Y.)2026

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

Conner J Langeberg, Taehan Kim, Roma Nagle, Agni Rajinikanth, Charlotte Meredith, Dimple A Garuadapuri, Jennifer Doudna, Jamie Cate

Abstract read
PubMed Publisher
In one paragraph

Article in RNA (New York, N.Y.), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

8 authors.

Conner J LangebergUniversity of California Berkeley.ORCID http://orcid.org/0000-0002-5609-3758
Taehan KimUniversity of California Berkeley.
Roma NagleUniversity of California Berkeley.
Agni RajinikanthUniversity of California Berkeley.
Charlotte MeredithUniversity of California Berkeley.
Dimple A GaruadapuriUniversity of California Berkeley.
Jennifer DoudnaUniversity of California Berkeley.ORCID http://orcid.org/0000-0001-9161-999X
Jamie CateUniversity of California Berkeley j-h-doudna-cate@berkeley.edu.ORCID http://orcid.org/0000-0001-5965-7902

Funding

Project 3U54AI170792 · NIAID · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI Nevan J Krogan · 2022 to 2026
$35.7M
Resource Core II: In Vivo CoreU19NS132303 · NINDS · UNIVERSITY OF CALIFORNIA BERKELEY · PI NIREN MURTHY · 2023 to 2026
$19.7M
Mechanisms of Translation Control in HumansR35GM148352 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI JAMIE H CATE · 2023 to 2026
$2.3M
Expanding CRISPR-Cas editing technology through exploration of novel Cas proteins and DNA repair systemsU01AI142817 · NIAID · UNIVERSITY OF CALIFORNIA BERKELEY · PI BANFIELD, JILLIAN, DOUDNA, JENNIFER A · 2018 to 2022
$2.0M
Cas9 RNP delivery to immune cells in vivo via molecular targetingUH3AI150552 · NIAID · UNIVERSITY OF CALIFORNIA BERKELEY · PI DOUDNA, JENNIFER A, WILSON, ROSS C · 2022 to 2022
$1.3M
NIAID NIH HHS U01 AI142817NIAID NIH HHS U54 AI170792NIAID NIH HHS UH3 AI150552NIGMS NIH HHS R35 GM148352NINDS NIH HHS U19 NS132303
6 · The paper itself

Abstract

In recent years, deep-learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown moderate success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assessed the utility of this enhanced dataset by retraining on a deep-learning model, SincFold. We find that SincFold exhibited improved performance on a set of previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. We additionally implemented Lyra-TransPred, which achieved the highest mean F1 and MCC among the evaluated models while requiring substantially less training time per epoch. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Indexed as

Deep learningRfamRNA secondary structureRNASSTR

Identifiers

PMID42680564

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.