Evidence map›Paper›PMID 40016653›Full record

ArticleBMC bioinformatics2025

Comparative Assessment of Protein Large Language Models for Enzyme Commission Number Prediction.

João Capela, Maria Zimmermann-Kogadeeva, Aalt D J van Dijk, Dick de Ridder, Oscar Dias, Miguel Rocha

Abstract readComparative Study
In one paragraph

Article in BMC bioinformatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed.

  1. Review
  2. Review
  3. Article
  4. Article
  5. Article
  6. Article
  7. Emerging technologies and current challenges in intratumoral microbiota research.Frontiers in cellular and infection microbiology · 2025
    Review
  8. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

João CapelaCentre of Biological Engineering, University of Minho, Braga, 4710-057, Portugal. joao.capela@ceb.uminho.pt.
Maria Zimmermann-KogadeevaGenome Biology Unit, European Molecular Biology Laboratory, Heidelberg, Germany.
Aalt D J van DijkBioinformatics Group, Department of Plant Sciences, Wageningen University and Research, Wageningen, The Netherlands.
Dick de RidderBioinformatics Group, Department of Plant Sciences, Wageningen University and Research, Wageningen, The Netherlands.
Oscar DiasCentre of Biological Engineering, University of Minho, Braga, 4710-057, Portugal.
Miguel RochaCentre of Biological Engineering, University of Minho, Braga, 4710-057, Portugal.

Funding

Fundação para a Ciência e a Tecnologia 10.54499/CEECIND/03425/2018/CP1581/CT0020Fundação para a Ciência e a Tecnologia DFA/BD/08789/2021LABBELS LA/P/0029/2020
6 · The paper itself

Abstract

backgroundProtein large language models (LLM) have been used to extract representations of enzyme sequences to predict their function, which is encoded by enzyme commission (EC) numbers. However, a comprehensive comparison of different LLMs for this task is still lacking, leaving questions about their relative performance. Moreover, protein sequence alignments (e.g. BLASTp or DIAMOND) are often combined with machine learning models to assign EC numbers from homologous enzymes, thus compensating for the shortcomings of these models' predictions. In this context, LLMs and sequence alignment methods have not been extensively compared as individual predictors, raising unaddressed questions about LLMs' performance and limitations relative to the alignment methods. In this study, we set out to assess the performance of ESM2, ESM1b, and ProtBERT language models in their ability to predict EC numbers, comparing them with BLASTp, against each other and against models that rely on one-hot encodings of amino acid sequences.

resultsOur findings reveal that combining these LLMs with fully connected neural networks surpasses the performance of deep learning models that rely on one-hot encodings. Moreover, although BLASTp provided marginally better results overall, DL models provide results that complement BLASTp's, revealing that LLMs better predict certain EC numbers while BLASTp excels in predicting others. The ESM2 stood out as the best model among the LLMs tested, providing more accurate predictions on difficult annotation tasks and for enzymes without homologs.

conclusionsCrucially, this study demonstrates that LLMs still have to be improved to become the gold standard tool over BLASTp in mainstream enzyme annotation routines. On the other hand, LLMs can provide good predictions for more difficult-to-annotate enzymes, particularly when the identity between the query sequence and the reference database falls below 25%. Our results reinforce the claim that BLASTp and LLM models complement each other and can be more effective when used together.

Indexed as

EnzymesProteinsComputational BiologyDatabases, ProteinLarge Language ModelsMachine LearningNeural Networks, ComputerSequence AlignmentEnzymesProteinsDeep learningEnzyme annotationLarge language models

Identifiers

PMID40016653
PMCPMC11866580

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.