ArticleScientific reports2025
Medium-sized protein language models perform well at transfer learning on realistic datasets.
Article in Scientific reports, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 18 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
18 citing papers in PubMed.
- From sites to structure to serology: a roadmap for structure-aware molecular evolution of antigenically evolving viruses.Journal of virology · 2026Review
- Engineering selective amyloid precursor protein inhibitors by machine learning and deep mutational scanning.Protein science : a publication of the Protein Society · 2026Article
- An enzyme-specific protein language model for catalytic property prediction.Nature communications · 2026Article
- GeoPep: A Geometry-Aware Masked Language Model for Protein-Peptide Binding Site Prediction.Journal of chemical information and modeling · 2026Article
- AbTune: layer-wise selective fine-tuning of protein language models for antibodies.Briefings in bioinformatics · 2026Article
- PUFFIN: protein unit discovery with functional supervision.Bioinformatics (Oxford, England) · 2026Article
- Machine learning framework for cost effective deep mutational scanning through targeted substitution profiling.BMC bioinformatics · 2026Article
- Article
- Discovering naturally occurring antifreeze peptides from microbiome by integrating protein language models and molecular dynamics simulation.Journal of materials chemistry. B · 2026Article
- Intrinsic dataset features drive mutational effect prediction by protein language models.bioRxiv : the preprint server for biology · 2026Article
- Protein language models enable accurate viral host range prediction.Scientific reports · 2026Article
- Learning physical interactions to compose biological large language models.Communications chemistry · 2026Review
- Architectural good practices for reproducible benchmarking in protein machine learning.Frontiers in bioinformatics · 2026Article
- Evaluating Pretrained Protein Language Model Embeddings as Proxies for Functional Similarity.Journal of molecular evolution · 2025Article
- A deep learning framework for lysine 2-hydroxyisobutyrylation site prediction using evolutionary feature representation.Scientific reports · 2025Article
- Low-N Protein Activity Optimization with FolDE.ArXiv · 2025Article
- Mechanistic modeling or machine learning for detecting variants of concern: Why not both?Proceedings of the National Academy of Sciences of the United States of America · 2025Article
- Without safeguards, AI-Biology integration risks accelerating future pandemics.Frontiers in microbiology · 2025Article
Corrections and comments
- Update of
Authors and funding
3 authors.
Funding
Abstract
Protein language models (pLMs) can offer deep insights into evolutionary and structural properties of proteins. While larger models, such as the 15 billion parameter model ESM-2, promise to capture more complex patterns in sequence space, they also present practical challenges due to their high dimensionality and high computational cost. We systematically evaluated the performance of various ESM-style models across multiple biological datasets to assess the impact of model size on transfer learning via feature extraction. Surprisingly, we found that larger models do not necessarily outperform smaller ones, in particular when data is limited. Medium-sized models, such as ESM-2 650M and ESM C 600M, demonstrated consistently good performance, falling only slightly behind their larger counterparts-ESM-2 15B and ESM C 6B-despite being many times smaller. Additionally, we compared various methods of compressing embeddings prior to transfer learning, and we found that mean embeddings consistently outperformed other compression methods. In summary, ESM C 600M with mean embeddings offers an optimal balance between performance and efficiency, making it a practical and scalable choice for transfer learning in realistic biological applications.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.