Evidence map›Paper›PMID 41739924›Full record

ArticleScience advances2026

Diverse database and machine learning model to narrow the generalization gap in RNA structure prediction.

Albéric A de Lajarte, Yves J Martin des Taillades, Justin Aruda, Pierre Bongrand, Federico Fuchs Wightman, Dragui Salazar, Matthew F Allan, Colin Kalicki, Casper L'Esperance-Kerckhoff, Alex Kashi and 2 more

Abstract read
In one paragraph

Article in Science advances, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Review
  5. Article
  6. Article
  7. Assessment of nucleic acid structure prediction in CASP16.bioRxiv : the preprint server for biology · 2025
    Article
  8. Article
  9. Review
  10. Transformers in RNA structure prediction: A review.Computational and structural biotechnology journal · 2025
    Review
  11. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Albéric A de LajarteDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0009-0003-2637-5820
Yves J Martin des TailladesDepartment of Biochemistry, Stanford University, Stanford, CA, USA.ORCID 0000-0003-4954-612X
Justin ArudaDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0000-0001-7526-6135
Pierre BongrandDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0000-0002-7228-1898
Federico Fuchs WightmanDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0000-0003-3934-3628
Dragui SalazarDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0009-0004-9548-2555
Matthew F AllanDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0000-0001-8182-7402
Colin KalickiDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.
Casper L'Esperance-KerckhoffDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0009-0002-7086-7323
Alex KashiDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.
Fabrice JossinetFaculty of Life Sciences, University of Strasbourg, Strasbourg, France.ORCID 0000-0003-4384-5735
Silvi RouskinDepartment of Microbiology, Harvard Medical School, Boston, MA, USA.ORCID 0000-0003-2042-6642

Funding

Constructing the nest - understanding the mechanisms of nidoviridae RNA genomes transcription and recombinationDP2AI175475 · NIAID · HARVARD MEDICAL SCHOOL · PI ROUSKIN, SILVIA · 2022 to 2022
$1.5M
NIAID NIH HHS DP2 AI175475Wellcome Trust
6 · The paper itself

Abstract

Understanding macromolecular structures of proteins and nucleic acids is critical for discerning their functions and biological roles. Advanced techniques-crystallography, nuclear magnetic resonance, and cryo-electron microscopy-have facilitated the determination of more than 180,000 protein structures, all cataloged in the Protein Data Bank. This comprehensive repository has been pivotal in developing deep learning algorithms for predicting protein structures directly from sequences. In contrast, RNA structure prediction has lagged and suffers from a scarcity of structural data. Here, we present the secondary structure models of 1098 primary microRNAs and 1456 human messenger RNA regions determined through chemical probing. We develop a deep learning architecture inspired from the Evoformer model of Alphafold and traditional architectures for secondary structure prediction. This model, eFold, was trained on our newly generated database and more than 300,000 secondary structures across multiple sources. We benchmark eFold on two challenging test sets of long and diverse RNA structures and show that our dataset and architecture contribute to increasing the prediction performance, compared to similar state-of-the-art methods. Together, our results reveal that merely expanding the database size is insufficient for generalization across families, whereas incorporating a greater diversity and complexity of RNA structures allows for enhanced model performance.

Indexed as

Computational BiologyMachine LearningNucleic Acid ConformationRNARNA, MessengerAlgorithmsHumansMicroRNAsModels, MolecularPrediction AlgorithmsMicroRNAsRNARNA, Messenger

Identifiers

PMID41739924
PMCPMC12935039

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.