Evidence map›Paper›PMID 38653796›Full record

ArticleNature biotechnology2025

Computational scoring and experimental evaluation of enzymes generated by neural networks.

Sean R Johnson, Xiaozhi Fu, Sandra Viknander, Clara Goldin, Sarah Monaco, Aleksej Zelezniak, Kevin K Yang

Erratum issuedOpen access · hybridAbstract read
In one paragraph

Article in Nature biotechnology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 45 papers.

0numbers the graph read from it
0cells of the map it votes in
45citing papers in PubMed
14.7field-weighted citation impact, top 1% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

45 citing papers in PubMed, 64 citations in OpenAlex.

  1. Review
  2. Article
  3. Review
  4. Article
  5. Zero-shot design of abioRxiv : the preprint server for biology · 2026
    Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. Article
  11. Review
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. A protein dynamics-based deep learning model enhances predictions of fitness and epistasis.Proceedings of the National Academy of Sciences of the United States of America · 2025
    Article
  18. GeoEvoBuilder: A deep learning framework for efficient functional and thermostable protein design.Proceedings of the National Academy of Sciences of the United States of America · 2025
    Article
  19. Before LUCA: unearthing the chemical roots of metabolism.Philosophical transactions of the Royal Society of London. Series B, Biological sciences · 2025
    Article
  20. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

7 authors at 5 institutions in 4 countries.

Sean R Johnson *New England Biolabs, Ipswich, MA, USA.ORCID http://orcid.org/0000-0001-8261-9015
Xiaozhi Fu *Department of Life Sciences, Chalmers University of Technology, Gothenburg, Sweden.ORCID http://orcid.org/0000-0002-4465-7528
Sandra ViknanderDepartment of Life Sciences, Chalmers University of Technology, Gothenburg, Sweden.
Clara GoldinDepartment of Life Sciences, Chalmers University of Technology, Gothenburg, Sweden.ORCID http://orcid.org/0009-0006-8235-3425
Aleksej ZelezniakDepartment of Life Sciences, Chalmers University of Technology, Gothenburg, Sweden. aleksej.zelezniak@chalmers.se.ORCID http://orcid.org/0000-0002-3098-9441
Kevin K YangMicrosoft Research, Cambridge, MA, USA. yang.kevin@microsoft.com.ORCID http://orcid.org/0000-0001-9045-6826
Chalmers University of Technology · SEInvitae (United States) · USMicrosoft (United States) · USNew England Biolabs (United States) · USVilnius University · LT

Funding

Svenska Forskningsrådet Formas (Swedish Research Council Formas) 2019-01403Vetenskapsrådet (Swedish Research Council) 2019-05356Vetenskapsrådet (Swedish Research Council) 2022-06725
6 · The paper itself

Abstract

In recent years, generative protein sequence models have been developed to sample novel sequences. However, predicting whether generated proteins will fold and function remains challenging. We evaluate a set of 20 diverse computational metrics to assess the quality of enzyme sequences produced by three contrasting generative models: ancestral sequence reconstruction, a generative adversarial network and a protein language model. Focusing on two enzyme families, we expressed and purified over 500 natural and generated sequences with 70-90% identity to the most similar natural sequences to benchmark computational metrics for predicting in vitro enzyme activity. Over three rounds of experiments, we developed a computational filter that improved the rate of experimental success by 50-150%. The proposed metrics and models will drive protein engineering research by serving as a benchmark for generative protein sequence models and helping to select active variants for experimental testing.

Indexed as

Computational BiologyEnzymesNeural Networks, ComputerAlgorithmsAmino Acid SequenceProtein EngineeringEnzymes

Identifiers

PMID38653796
PMCPMC11919684
OpenAlexW4395048825

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.