ArticleScientific reports2026
Machine learning based approaches for structure activity relationship analysis of heparanase inhibitors.
Article in Scientific reports, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Human Heparanase (HPSE), the only mammalian endo-β-D-glucuronidase, plays an important role in extracellular matrix remodeling and the release of heparin-bound growth factors. Its overexpression is strongly correlated with increased tumor growth, angiogenesis, metastasis, and inflammation, highlighting HPSE as a compelling therapeutic target for oncology and inflammatory diseases. This study aimed to develop and validate a robust computational workflow for predicting the activity class of potential HPSE inhibitors using curated data from the ChEMBL database. Bioactivity data ([Formula: see text]) for known HPSE inhibitors were extracted and put through a meticulous data curation process, which included chemical structure standardization, molecular weight filtering, and final deduplication based on standardized isomeric SMILES to ensure structural uniqueness. Continuous [Formula: see text] values (nM) were converted to [Formula: see text] and subsequently categorized into three activity classes: A ([Formula: see text] and [Formula: see text]), B ([Formula: see text] and [Formula: see text]), and C ([Formula: see text] and [Formula: see text]) for multi-class classification. Molecular representations included two-dimensional physicochemical descriptors, Morgan fingerprints, and three-dimensional descriptors derived from optimized low-energy conformers generated using ETKDGv3 and MMFF94s. Multiple machine learning classifiers were evaluated using pipelines incorporating imputation, scaling, optional Principal Component Analysis (PCA) dimensionality reduction applied to the combined feature sets, and SMOTE (Synthetic Minority Over-sampling Technique) to address class imbalance. Models were trained and optimized using randomized search cross-validation on an 80% training split, maximizing balanced accuracy. The best-performing model pipeline (RF_B, a Random Forest with PCA on 2D+Morgan Fingerprints+3D features) achieved approximately 80% accuracy and 78.5% balanced accuracy on the held-out 20% test set. The final validated model was successfully utilized to predict the activity classes of new, unseen compounds. This comprehensive pipeline provides a validated tool for classifying HPSE inhibitors derived from ChEMBL data, potentially aiding virtual screening efforts and guiding hit prioritization in drug discovery campaigns targeting HPSE.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.