Evidence map›Paper›PMID 40756675›Full record

ArticleBMC methods2025

ShortStop: a machine learning framework for microprotein discovery.

Brendan Miller, Eduardo Vieira de Souza, Victor J Pai, Hosung Kim, Joan M Vaughan, Calvin J Lau, Jolene K Diedrich, Alan Saghatelian

Abstract read
In one paragraph

Article in BMC methods, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Brendan MillerClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Eduardo Vieira de SouzaClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Victor J PaiClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Hosung KimUSC Stevens Neuroimaging and Informatics Institute, Keck School of Medicine of USC, University of Southern California, Los Angeles, CA USA.
Joan M VaughanClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Calvin J LauClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Jolene K DiedrichClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.
Alan SaghatelianClayton Foundation Laboratories for Peptide Biology, The Salk Institute for Biological Studies, 10010 N Torrey Pines Rd, San Diego, CA USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: The human genome contains over 3 million small open reading frames (smORFs, Methods: To address this challenge, we developed ShortStop, a computational framework based on the idea that not all translating smORFs produce functional proteins, but the ones that do may resemble experimentally characterized microproteins. ShortStop classifies smORFs into two reference groups: Swiss-Prot Analog Microproteins (SAMs), which resemble known microproteins, and PRISMs (Physicochemically Resembling In Silico Microproteins), which are synthetic sequences designed to match the composition of translating smORFs but lacking sequence order or evolutionary selection, and therefore serving as a proxy for non-functional peptides. This two-class system enables machine learning to help prioritize smORFs for downstream study. Results: ShortStop achieved high precision (90-94%), recall (87-96%), and F1 scores (90-93%) across all classes. When applied to a published dataset of translating smORFs, ShortStop classified about 8% as candidates with biochemical properties resembling Swiss-Prot microproteins (i.e., called SAMs). The remaining 92% resembled in silico generated sequences (i.e., called PRISMs), representing noncanonical proteins, non-functional peptides, or regulatory translation events. SAMs showed lower C-terminal hydrophobicity-linked to reduced proteasomal degradation-and greater N-terminal hydrophilicity at neutral pH, suggesting improved solubility and intracellular stability. ShortStop also identified microproteins overlooked by other methods, including one encoded by an upstream overlapping smORF in the StAR gene, which was detectable in human cells and steroid-producing tissues. In a clinical lung cancer dataset, ShortStop uncovered differentially expressed microprotein candidates, several of which were validated by mass spectrometry. Discussion: ShortStop addresses a key gap in microprotein research-the lack of scalable tools to characterize microproteins and standardized negative training data to train machine learning models for microproteins. By providing a classification framework rooted in biochemical features, ShortStop offers a practical solution for targeting smORFs in functional studies, benchmarking new discovery tools, and advancing microprotein research. Supplementary Information: The online version contains supplementary material available at 10.1186/s44330-025-00037-4.

Indexed as

CancerDe Novo genesMachine learningMicroproteinPeptidesProteogenomicsRibosome profilingSmall open reading frameSteroidogenic acute regulatory protein

Identifiers

PMID40756675
PMCPMC12313729

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.