Evidence map›Paper›PMID 40938959›Full record

ArticlePLoS computational biology2025

Paying attention to attention: High attention sites as indicators of protein family and function in language models.

Gowri Nayar, Alp Tartici, Russ B Altman

Abstract read
In one paragraph

Article in PLoS computational biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 9 papers.

0numbers the graph read from it
0cells of the map it votes in
9citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

9 citing papers in PubMed.

  1. PUFFIN: protein unit discovery with functional supervision.Bioinformatics (Oxford, England) · 2026
    Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Gowri NayarDepartment of Biomedical Data Science, Stanford University, Stanford, California, United States of America.ORCID 0000-0001-5819-7115
Alp TarticiDepartment of Genetics, Stanford University, Stanford, California, United States of America.
Russ B AltmanDepartment of Biomedical Data Science, Stanford University, Stanford, California, United States of America.ORCID 0000-0003-3859-2905

Funding

Computational methods for characterizing sources of variability in drug responseR35GM153195 · NIGMS · STANFORD UNIVERSITY · PI RUSS BIAGIO ALTMAN · 2024 to 2026
$1.0M
Exploring Understudied Proteins to Predict Novel Pathways and Associations to DiseaseF31LM014646 · NLM · STANFORD UNIVERSITY · PI NAYAR, GOWRI · 2024 to 2025
$92k
NIGMS NIH HHS R35 GM153195NLM NIH HHS F31 LM014646
6 · The paper itself

Abstract

Protein Language Models (PLMs) use transformer architectures to capture patterns within protein primary sequences, providing a powerful computational representation of the amino acid sequence. Through large-scale training on protein primary sequences, PLMs generate vector representations that encapsulate the biochemical and structural properties of proteins. At the core of PLMs is the attention mechanism, which facilitates the capture of long-range dependencies by computing pairwise importance scores across residues, thereby highlighting regions of biological interaction within the sequence. The attention matrices offer an untapped opportunity to uncover specific biological properties of proteins, particularly their functions. In this work, we introduce a novel approach, using the Evolutionary Scale Modelling (ESM), for identifying High Attention (HA) sites within protein primary sequences, corresponding to key residues that define protein families. By examining attention patterns across multiple layers, we pinpoint residues that contribute most to family classification and function prediction. Our contributions are as follows: (1) we propose a method for identifying HA sites at critical residues from the middle layers of the PLM; (2) we demonstrate that these HA sites provide interpretable links to biological functions; and (3) we show that HA sites improve active site predictions for functions of unannotated proteins. We make available the HA sites for the human proteome. This work offers a broadly applicable approach to protein classification and functional annotation and provides a biological interpretation of the PLM's representation.

Indexed as

ProteinsAlgorithmsAmino Acid SequenceComputational BiologyDatabases, ProteinHumansSequence Analysis, ProteinProteins

Identifiers

PMID40938959
PMCPMC12448987

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.