Evidence map›Paper›PMID 42298320›Full record

ArticleJournal of the American Society for Mass Spectrometry2026

Benchmarking MS/MS Featurization Strategies for Machine Learning-Driven Metabolite Structure Annotation.

Roger Giné, Ivan Pérez-López, Josep M Badia, Jordi Capellades, Oscar Yanes

Abstract read
In one paragraph

Article in Journal of the American Society for Mass Spectrometry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Roger GinéUniversitat Rovira i Virgili, Department of Electronic Engineering, 43007 Tarragona, Spain.ORCID 0000-0003-0288-9619
Ivan Pérez-LópezUniversitat Rovira i Virgili, Department of Electronic Engineering, 43007 Tarragona, Spain.
Josep M BadiaUniversitat Rovira i Virgili, Department of Electronic Engineering, 43007 Tarragona, Spain.
Jordi CapelladesCIBER de Diabetes y Enfermedades Metabólicas Asociadas (CIBERDEM), Instituto de Salud Carlos III, 28029 Madrid, Spain.
Oscar YanesUniversitat Rovira i Virgili, Department of Electronic Engineering, 43007 Tarragona, Spain.ORCID 0000-0003-3695-7157

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Reference MS/MS libraries remain incomplete due to the vast chemical diversity of metabolites, leaving many spectra from untargeted metabolomics experiments unannotated─the "dark matter" of metabolomics. Machine learning can extend metabolite annotation beyond direct library matches, but its success depends critically on how MS/MS spectra are converted into numerical representations that capture chemically meaningful features while reducing sparsity. Although numerous spectral representations exist, they have not been systematically compared. Using over 71,000 unique compounds with merged-energy MS/MS spectra, we benchmarked a broad set of spectral featurization methods, including fixed and adaptive binning, global-quantile variable-width bins, frequent-peaks representations, spectrum hashing, and learned embeddings such as Spec2Vec, MS2DeepScore, DreaMS, and SpecEmbedding. We further evaluated how vector dimensionality affects performance. A total of 105 neural network models were trained under 5-fold cross-validation to predict Mol2Vec molecular embeddings and retrieve correct structures from a 0.6-million-compound database. Retrieval was assessed at 0.1, 3, and 10 ppm mass tolerances, and a null ranking model was generated to determine expected Top-N accuracy under random candidate ordering. Adaptive binning, frequent-peaks, and DreaMS produced the most accurate embedding predictions. On the test data set, Top-1 retrieval reached 46%, 44%, and 38% for 0.1, 3, and 10 ppm, respectively, with Top-5 accuracies up to 77%. In the CASMI2022 data set, Top-1 performance remained similar at 0.1 ppm but dropped markedly at wider tolerances, reaching only 26% at 3 ppm and 23% at 10 ppm. To ensure reproducibility and broad community applicability, results were further validated on two fully open benchmark data sets, MassSpecGym and Spectraverse, with findings consistent across all three resources. These results underscore clear performance differences among featurization strategies, the strong dependence of retrieval accuracy on mass precision, and the need for evaluation metrics aligned with structure-level annotation tasks.

Indexed as

BenchmarkingMetabolomicsMachine LearningTandem Mass Spectrometrymachine learningmetabolite annotationspectral featurizationstructure retrievaltandem mass spectrometryuntargeted metabolomics

Identifiers

PMID42298320
PMCPMC13329996

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.