Evidence map›Paper›PMID 42292907›Full record

ArticlebioRxiv : the preprint server for biology2026

Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations.

Phillip R Woolley, Aaron L Feller, Andrew D Ellington, Claus O Wilke

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Phillip R WoolleyDepartment of Molecular Biosciences, The University of Texas at Austin.ORCID 0000-0003-4473-9424
Aaron L FellerDepartment of Molecular Biosciences, The University of Texas at Austin.ORCID 0000-0002-4476-1026
Andrew D EllingtonDepartment of Molecular Biosciences, The University of Texas at Austin.ORCID 0000-0001-6246-5338
Claus O WilkeDepartment of Integrative Biology, The University of Texas at Austin.ORCID 0000-0002-7470-9261

Funding

Directed evolution of broadly fungible biosensorsR01GM146093 · NIGMS · UNIVERSITY OF TEXAS AT AUSTIN · PI Andrew D Ellington · 2023 to 2026
$1.3M
NIGMS NIH HHS R01 GM146093
6 · The paper itself

Abstract

Deep learning models have emerged as promising tools in protein engineering. In particular, they can be used to predict mutation fitness without the need for task-specific training, a process known as zero-shot prediction. However, the respective strengths and limitations of zero-shot predictions remain poorly understood. Here, we argue that commonly used large-scale benchmarks obscure important failure modes relevant to practical protein engineering, including an inability to pinpoint highly fit mutations or variants driving new-to-nature functions. Beyond these practical failures, we identify a fundamental limitation of zero-shot prediction: a generic fitness score cannot simultaneously optimize for distinct, competing engineering targets, meaning it is inherently disconnected from the phenotype of interest. Moreover, in a systematic comparison of a wide range of available models, we demonstrate that most models show comparable zero-shot performance, irrespective of model architecture and/or input modality (sequence vs. structure). Ultimately, we find that zero-shot predictions serve only as coarse filters separating fit mutations from deleterious ones, failing to reliably identify the mutations that would be most valuable in protein engineering.

Indexed as

Deep LearningMutation PredictionProtein Language ModelZero-shot

Identifiers

PMID42292907
PMCPMC13261798

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.