Evidence map›Paper›PMID 42507936›Full record

ArticleProceedings of the National Academy of Sciences of the United States of America2026

Discovery of a phenazine-thiol conjugase from sparse data using genome-informed machine learning.

Xiaoyu Shan, Inês B Trindade, Nathaniel R Glasser, Korbinian O Thalhammer, Matthew Scurria, Ariane Mora, Stuart J Conway, Dianne K Newman

Abstract read
In one paragraph

Article in Proceedings of the National Academy of Sciences of the United States of America, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

8 authors.

Xiaoyu Shan *Division of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA 91125.ORCID 0000-0001-9631-3244
Inês B Trindade *Division of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA 91125.
Nathaniel R GlasserResnick Sustainability Institute, California Institute of Technology, Pasadena, CA 91125.ORCID 0000-0002-2833-5166
Korbinian O ThalhammerDivision of Geological and Planetary Sciences, California Institute of Technology, Pasadena, CA 91125.
Matthew ScurriaDepartment of Chemistry and Biochemistry, University of California, Los Angeles, CA 90095.
Ariane MoraAITHYRA GmbH, Research Institute for Biomedical Artificial Intelligence of the Austrian Academy of Sciences, Vienna 1030, Austria.ORCID 0000-0003-1331-8192
Stuart J ConwayDepartment of Chemistry and Biochemistry, University of California, Los Angeles, CA 90095.
Dianne K NewmanDivision of Biology and Biological Engineering, California Institute of Technology, Pasadena, CA 91125.ORCID 0000-0003-1647-1918

Funding

Biological mechanisms and consequences of efficient extracellular electron transfer in Pseudomonas aeruginosaR01AI127850 · NIAID · CALIFORNIA INSTITUTE OF TECHNOLOGY · PI Dianne K Newman · 2017 to 2026
$5.6M
European Molecular Biology Organization (EMBO) ALTF 191-2023HHS | NIH | National Institute of Allergy and Infectious Diseases (NIAID) 2R01AI127850-06A1)Jung Chair in Medicinal Chemistry and Drug Discovery UCLA NANemko Postdoctoral Fellowship Caltech BBE Division NANIAID NIH HHS R01 AI127850Schmidt Science/Rhodes Trust NA
6 · The paper itself

Abstract

Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (phenazine-thiol conjugase), an enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through nonenzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.

Indexed as

Machine LearningPhenazinesGlutathioneGlutathionephenazinePhenazinescomparative genomicsenzyme discoveryglutathionemachine learningphenazine

Identifiers

PMID42507936
PMCPMC13419639

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.