Evidence map›Paper›PMID 42635244›Full record

ArticleBioinformatics (Oxford, England)2026

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

Georg Auge, Matthis Clausen, Konstantin Ketterer, Jacob Schaefer, Nils Schmitt, Tom Altenburg, Yannick Hartmaring, Hendrik Raetz, Christoph N Schlaffner, Bernhard Y Renard

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Georg AugeHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0005-4234-1668
Matthis ClausenHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0001-5548-7595
Konstantin KettererHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0005-9816-4306
Jacob SchaeferHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0008-4979-9833
Nils SchmittHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0008-9971-3727
Tom AltenburgHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0000-0002-6435-4073
Yannick HartmaringHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0009-0005-0043-7709
Hendrik RaetzHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0000-0002-1230-1236
Christoph N SchlaffnerHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0000-0003-2717-3406
Bernhard Y RenardHasso Plattner Institute, Digital Engineering Faculty, University of Potsdam, 14482 Potsdam, Germany.ORCID 0000-0003-4589-9809

Funding

European Research Council 101124385European Research Council eXplAInProt
6 · The paper itself

Abstract

motivationAn unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale.

resultsWithin 49 h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Indexed as

Databases, ProteinData CurationMachine LearningProteomicsSoftwareMass Spectrometry

Identifiers

PMID42635244
PMCPMC13501284

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.