Evidence map›Paper›PMID 40166178›Full record

ArticlebioRxiv : the preprint server for biology2025

Pool PaRTI: A PageRank-Based Pooling Method for Identifying Critical Residues and Enhancing Protein Sequence Representations.

Alp Tartici, Gowri Nayar, Russ B Altman

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

3 authors.

Alp TarticiStanford University.ORCID 0009-0008-1885-0077
Gowri NayarStanford University.ORCID 0000-0001-5819-7115
Russ B AltmanStanford University.ORCID 0000-0003-3859-2905

Funding

Undergraduate Summer Research Experiences Support for Combining systems biology and structural biology to find new therapeuticsR01GM102365 · NIGMS · STANFORD UNIVERSITY · PI ALTMAN, RUSS BIAGIO · 2012 to 2021
$3.0M
Computational methods for characterizing sources of variability in drug responseR35GM153195 · NIGMS · STANFORD UNIVERSITY · PI RUSS BIAGIO ALTMAN · 2024 to 2026
$1.0M
Exploring Understudied Proteins to Predict Novel Pathways and Associations to DiseaseF31LM014646 · NLM · STANFORD UNIVERSITY · PI NAYAR, GOWRI · 2024 to 2025
$92k
NIGMS NIH HHS R01 GM102365NIGMS NIH HHS R35 GM153195NLM NIH HHS F31 LM014646
6 · The paper itself

Abstract

Motivation: Protein language models produce token-level embeddings for each residue, resulting in an output matrix with dimensions that vary based on sequence length. However, downstream machine learning models typically require fixed-length input vectors, necessitating a pooling method to compress the output matrix into a single vector representation of the entire protein. Traditional pooling methods often result in substantial information loss, impacting downstream task performance. We aim to develop a pooling method that produces more expressive general-purpose protein embedding vectors while offering biological interpretability. Results: We introduce Pool PaRTI, a novel pooling method that leverages internal transformer attention matrices and PageRank to assign token importance weights. Our unsupervised and parameter-free approach consistently prioritizes residues experimentally annotated as critical for function, assigning them higher importance scores. Across four diverse protein machine learning tasks, Pool PaRTI enables significant performance gains in predictive performance. Additionally, it enhances interpretability by identifying biologically relevant regions without relying on explicit structural data or annotated training. To assess generalizability, we evaluated Pool PaRTI with two encoder-only protein language models, confirming its robustness across different models. Availability and Implementation: Pool PaRTI is implemented in Python with PyTorch and is available at https://github.com/Helix-Research-Lab/Pool_PaRTI.git. The Pool PaRTI sequence embeddings and residue importance values for all human proteins on UniProt are available at https://zenodo.org/records/15036725 for ESM2 and protBERT.

Identifiers

PMID40166178
PMCPMC11956911

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.