ArticleFrontiers in molecular biosciences2026
Geographic authentication of premium tobacco extracts using machine learning models trained on LC-HRMS fingerprints.
Article in Frontiers in molecular biosciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
12 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Introduction: The geographic origin of premium tobacco extracts fundamentally shapes their sensory profiles and commercial value, but authenticating these complex matrices remains a formidable analytical challenge. Conventional nontargeted analysis workflows typically rely on a minute fraction of structurally annotated metabolites, thereby discarding the vast majority of unannotated yet highly specific chemical markers. To address this limitation, this study introduces a geographic authentication framework utilizing machine learning models trained on characteristic ion feature (CIF) fingerprints. Methods: A large-scale dataset of 1,697 tobacco samples was analyzed using liquid chromatography-high resolution mass spectrometry coupled with a data-independent acquisition strategy. Rigorous statistical filtering was applied to isolate CIFs and strip away the shared background metabolome. To evaluate the samples, univariate thresholding based on pattern similarity (dot product) and marker intensity ratios (R score) was initially tested. To overcome the vulnerability of univariate methods to batch effects and overlapping chemical profiles, optimized multivariate machine learning models were also deployed. The models were evaluated under a strict Leave-One-Batch-Out cross-validation framework to ensure robust cross-batch generalization. Results: Statistical filtering effectively stripped away over 90% of the shared background metabolome, isolating 47,543 CIFs for Zimbabwe tobacco and 3,646 CIFs for Yunnan tobacco. Univariate thresholding achieved a maximum balanced accuracy of 77.9% but proved too rigid. Under cross-validation, the multivariate Random Forest (RF) classifier demonstrated strong predictive performance, achieving a balanced accuracy of 74.6% for Zimbabwe tobacco and successfully detecting economically motivated adulteration across a wide dilution range (0-1,000 ppm). Conversely, both univariate and multivariate methods failed to achieve robust authentication for the Yunnan extracts. This underperformance was driven by severe intra-group heterogeneity stemming from diverse suppliers and heterogeneous processing treatments, which heavily diluted universal characteristic features. Discussion: Ultimately, this hybrid strategy provides a scalable tool for the industrial verification of consistent premium extracts. Furthermore, the challenges encountered with the Yunnan samples underscore the necessity of cohort-specific modeling when dealing with highly diverse agricultural products that are subjected to varied processing treatments.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.