Evidence map›Paper›PMID 38187750›Full record

ArticlebioRxiv : the preprint server for biology2025

Unexplored regions of the protein sequence-structure map revealed at scale by a library of foldtuned language models.

Arjuna M Subramanian, Zachary A Martinez, Alec L Lourenço, Sonia C Yuan, Shichen Liu, Matt Thomson

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Arjuna M SubramanianDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.ORCID 0009-0004-2790-0209
Zachary A MartinezDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.ORCID 0000-0002-7830-3162
Alec L LourençoDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.
Sonia C YuanDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.ORCID 0009-0004-0257-9818
Shichen LiuDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.ORCID 0000-0002-0964-6559
Matt ThomsonDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA.

Funding

A Global Map of Interactions Among Human Cell Surface Proteins and Secreted LigandsR01GM150125 · NIGMS · CALIFORNIA INSTITUTE OF TECHNOLOGY · PI GARCIA, KENAN CHRISTOPHER, THOMSON, MATTHEW W. · 2022 to 2025
$11.3M
NIGMS NIH HHS R01 GM150125
6 · The paper itself

Abstract

The combinatorial scale of amino-acid sequence-space has traditionally precluded substantive study of the full protein sequence-structure map. It remains unknown, for instance, how much of the vast uncharted landscape of far-from-natural sequences encodes the familiar ensemble of natural folds in a fashion consistent with the laws of biophysics but seemingly untouched by evolution on Earth. The scale of sequence perturbations required to access these spaces exceeds the reach of even gold-standard experimental approaches such as directed evolution. We surpass this limitation guided by the innate capacity of protein language models (PLMs) to explore sequences outside their natural training data through generation and self-feedback. We recast PLMs as probes that explore into regions of protein "deep space" that possess little-to-no detectable homology to natural examples, while enforcing core structural constraints, in a novel sequence design approach that we term "foldtuning." We build a library of foldtuned PLMs for >700 natural folds in the SCOP database, covering numerous high-priority targets for synthetic biology, including GPCRs and small GTPases, composable cell-surface-receptor and DNA-binding domains, and small signaling/regulatory domains. Candidate proteins generated by foldtuned PLMs reflect distinctive new "rules of language" for sequence innovation beyond detectable homology to any known protein and sample subtle structural alterations in a manner reminiscent of natural structural evolution and diversification. Experimental validation of three markedly different fold targets; the tyrosine-kinase- and small-GTPase-regulating SH3 domain, the bacterial RNase inhibitor barstar, and the peptide hormone insulin demonstrates that foldtuning proposes protein variants that express and fold stably

Identifiers

PMID38187750
PMCPMC10769378

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.