Evidence map›Paper›PMID 38662579›Full record

ArticleBioinformatics (Oxford, England)2024

LMCrot: an enhanced protein crotonylation site predictor by leveraging an interpretable window-level embedding from a transformer-based protein language model.

Pawel Pratyush, Soufia Bahmani, Suresh Pokharel, Hamid D Ismail, Dukka B Kc

Open access · goldAbstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
4.2field-weighted citation impact, top 5% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed, 18 citations in OpenAlex.

  1. Predicting Enzyme Turnover Numbers and Enabling Rational Enzyme Evolution.Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2026
    Article
  2. Article
  3. Review
  4. Article
  5. PLM-eXplain: divide and conquer the protein embedding space.Bioinformatics (Oxford, England) · 2026
    Article
  6. Review
  7. Article
  8. Article
  9. Review
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors at 1 institution in 1 country.

Pawel PratyushDepartment of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.ORCID 0000-0002-4210-1200
Soufia BahmaniDepartment of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.
Suresh PokharelDepartment of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.
Hamid D IsmailDepartment of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.
Dukka B KcDepartment of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.ORCID 0000-0001-7443-1928
Michigan Technological University · US

Funding

National Science Foundation # 1901793
6 · The paper itself

Abstract

motivationRecent advancements in natural language processing have highlighted the effectiveness of global contextualized representations from protein language models (pLMs) in numerous downstream tasks. Nonetheless, strategies to encode the site-of-interest leveraging pLMs for per-residue prediction tasks, such as crotonylation (Kcr) prediction, remain largely uncharted.

resultsHerein, we adopt a range of approaches for utilizing pLMs by experimenting with different input sequence types (full-length protein sequence versus window sequence), assessing the implications of utilizing per-residue embedding of the site-of-interest as well as embeddings of window residues centered around it. Building upon these insights, we developed a novel residual ConvBiLSTM network designed to process window-level embeddings of the site-of-interest generated by the ProtT5-XL-UniRef50 pLM using full-length sequences as input. This model, termed T5ResConvBiLSTM, surpasses existing state-of-the-art Kcr predictors in performance across three diverse datasets. To validate our approach of utilizing full sequence-based window-level embeddings, we also delved into the interpretability of ProtT5-derived embedding tensors in two ways: firstly, by scrutinizing the attention weights obtained from the transformer's encoder block; and secondly, by computing SHAP values for these tensors, providing a model-agnostic interpretation of the prediction results. Additionally, we enhance the latent representation of ProtT5 by incorporating two additional local representations, one derived from amino acid properties and the other from supervised embedding layer, through an intermediate fusion stacked generalization approach, using an n-mer window sequence (or, peptide/fragment). The resultant stacked model, dubbed LMCrot, exhibits a more pronounced improvement in predictive performance across the tested datasets. AVAILABILITY AND IMPLEMENTATION: LMCrot is publicly available at https://github.com/KCLabMTU/LMCrot.

Indexed as

ProteinsAmino Acid SequenceComputational BiologyDatabases, ProteinNatural Language ProcessingProtein Processing, Post-TranslationalSoftwareProteins

Identifiers

PMID38662579
PMCPMC11088740
OpenAlexW4395118066

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.