Evidence map›Paper›PMID 42789669›Full record

ArticlePLoS computational biology2026

Fast structural search for classification of gut bacterial mucin O-glycan degrading enzymes.

Mert Erden, Tyler Schult, Karin Yanagi, Jugal Kishore Sahoo, David L Kaplan, Lenore J Cowen, Kyongbum Lee

Abstract read
In one paragraph

Article in PLoS computational biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Mert ErdenDepartment of Computer Science, Tufts University, Medford, Massachusetts, United States of America.
Tyler SchultDepartment of Chemical and Biological Engineering, Tufts University, Medford, Massachusetts, United States of America.ORCID https://orcid.org/0009-0009-3338-4219
Karin YanagiDepartment of Chemical and Biological Engineering, Tufts University, Medford, Massachusetts, United States of America.
Jugal Kishore SahooDepartment of Biomedical Engineering, Tufts University, Medford, Massachusetts, United States of America.
David L KaplanDepartment of Biomedical Engineering, Tufts University, Medford, Massachusetts, United States of America.
Lenore J CowenDepartment of Computer Science, Tufts University, Medford, Massachusetts, United States of America.ORCID https://orcid.org/0000-0001-6698-6413
Kyongbum LeeDepartment of Computer Science, Tufts University, Medford, Massachusetts, United States of America.ORCID https://orcid.org/0000-0002-0699-8057

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

The Enzyme Commission (EC) numbering scheme provides a hierarchical way to classify enzymes according to their catalytic functions. While recent protein language model (PLM) based approaches like CLEAN and ProteInter have improved sequence-based EC number prediction, they struggle with fine-grained classification at the deepest hierarchical level. Structure-based approaches for grouping similar proteins using alignment tools excel at finding proteins that share overall global structure, but suffer from high false positive rates when classifying proteins that are globally structurally similar but functional differentiation depends on a localized region. This problem is particularly relevant to EC number prediction, as enzymatic function depends on its catalytic domain, which is a relatively small, specific region of the protein. We introduce Deep Enzyme Function Transfer (DEFT) that harmonizes sequence- and structure-based approaches through the key insight that PLM based annotations of the first two EC number hierarchy levels vastly reduce false positives that are likely to show in purely structure-based EC number prediction. Given an enzyme of interest, DEFT first uses a PLM based method to assign the first two levels of the enzyme's EC number, and then uses a structure-based method to predict the remaining two levels of the EC number. Using benchmarking datasets, we demonstrate that DEFT achieves superior accuracy compared with current state-of-the-art tools for EC number prediction. Furthermore we show that DEFT's computational efficiency enables high-throughput, genome-wide annotations of total enzyme repertoires in organisms. We illustrate this capability by experimentally validating DEFT predicted glycoside hydrolase (GH) profiles of intestinal mucus associated bacteria.

Indexed as

Bacterial ProteinsMucinsPolysaccharidesComputational BiologyDatabases, ProteinGlycoside HydrolasesModels, MolecularBacterial ProteinsGlycoside HydrolasesMucinsPolysaccharides

Identifiers

PMID42789669
PMCPMC13626479

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.