ArticleBMC medical informatics and decision making2024
Machine learning pipeline to analyze clinical and proteomics data: experiences on a prostate cancer case.
Article in BMC medical informatics and decision making, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers, 1 of them a synthesis that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
8 citing papers in PubMed, 1 synthesis or guideline pooled it, 18 citations in OpenAlex.
- Progress and trends on machine learning in proteomics during 1997-2024: a bibliometric analysis.Frontiers in medicine · 2025Pooled it
- From Machine Learning-Enhanced Proteomics to a Validated Diagnostic Model: A Pipeline for Breast Cancer Biomarker Discovery via Independent and Transcriptomic Corroboration.Bioengineering (Basel, Switzerland) · 2026Article
- Article
- An ensemble-based model comprising deep learning for predicting peptide-binding residues in proteins.NAR genomics and bioinformatics · 2025Article
- A Comprehensive Proteome of Human Corneal Epithelial Cells Constructed by Cross-platform DIA-Mass Spectrometry.Scientific data · 2025Article
- Integrated Proteomics and Machine Learning Approach Reveals PYCR1 as a Novel Biomarker to Predict Prognosis of Sinonasal Squamous Cell Carcinoma.International journal of molecular sciences · 2024Article
- Optimization of diagnosis and treatment of hematological diseases via artificial intelligence.Frontiers in medicine · 2024Review
- Multi-omics based artificial intelligence for cancer research.Advances in cancer research · 2024Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors at 3 institutions in 1 country.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Proteomic-based analysis is used to identify biomarkers in blood samples and tissues. Data produced by devices such as mass spectrometry requires platforms to identify and quantify proteins (or peptides). Clinical information can be related to mass spectrometry data to identify diseases at an early stage. Machine learning techniques can be used to support physicians and biologists in studying and classifying pathologies. We present the application of machine learning techniques to define a pipeline aimed at studying and classifying proteomics data enriched using clinical information. The pipeline allows users to relate established blood biomarkers with clinical parameters and proteomics data. The proposed pipeline entails three main phases: (i) feature selection, (ii) models training, and (iii) models ensembling. We report the experience of applying such a pipeline to prostate-related diseases. Models have been trained on several biological datasets. We report experimental results about two datasets that result from the integration of clinical and mass spectrometry-based data in the contexts of serum and urine analysis. The pipeline receives input data from blood analytes, tissue samples, proteomic analysis, and urine biomarkers. It then trains different models for feature selection, classification and voting. The presented pipeline has been applied on two datasets obtained in a 2 years research project which aimed to extract hidden information from mass spectrometry, serum, and urine samples from hundreds of patients. We report results on analyzing prostate datasets serum with 143 samples, including 79 PCa and 84 BPH patients, and an urine dataset with 121 samples, including 67 PCa and 54 BPH patients. As results pipeline allowed to identify interesting peptides in the two datasets, 6 for the first one and 2 for the second one. The best model for both serum (AUC=0.87, Accuracy=0.83, F1=0.81, Sensitivity=0.84, Specificity=0.81) and urine (AUC=0.88, Accuracy=0.83, F1=0.83, Sensitivity=0.85, Specificity=0.80) datasets showed good predictive performances. We made the pipeline code available on GitHub and we are confident that it will be successfully adopted in similar clinical setups.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.