Evidence map›Paper›PMID 38584282›Full record

ArticleBMC medical informatics and decision making2024

Machine learning pipeline to analyze clinical and proteomics data: experiences on a prostate cancer case.

Patrizia Vizza, Federica Aracri, Pietro Hiram Guzzi, Marco Gaspari, Pierangelo Veltri, Giuseppe Tradigo

Open access · goldAbstract read
In one paragraph

Article in BMC medical informatics and decision making, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 8 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed, 1 pooled it
4.3field-weighted citation impact, top 5% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed, 1 synthesis or guideline pooled it, 18 citations in OpenAlex.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Review
  8. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors at 3 institutions in 1 country.

Patrizia VizzaDepartment of Surgical and Medical Sciences, Magna Græcia University, 88100, Catanzaro, Italy.
Federica AracriDepartment of Surgical and Medical Sciences, Magna Græcia University, 88100, Catanzaro, Italy. federica.aracri@unicz.it.
Pietro Hiram GuzziDepartment of Surgical and Medical Sciences, Magna Græcia University, 88100, Catanzaro, Italy.
Marco GaspariDepartment of Experimental and Clinical Medicine, Magna Græcia University, 88100, Catanzaro, Italy.
Pierangelo VeltriDepartment of Computers, Modeling, Electronics and Systems Engineering, University of Calabria, 87036, Rende, Italy.
Giuseppe TradigoDepartment of Theoretical and Applied Sciences, eCampus University, 22060, Novedrate, CO, Italy.
Magna Graecia University · ITUniversità degli Studi eCampus · ITUniversity of Calabria · IT

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Proteomic-based analysis is used to identify biomarkers in blood samples and tissues. Data produced by devices such as mass spectrometry requires platforms to identify and quantify proteins (or peptides). Clinical information can be related to mass spectrometry data to identify diseases at an early stage. Machine learning techniques can be used to support physicians and biologists in studying and classifying pathologies. We present the application of machine learning techniques to define a pipeline aimed at studying and classifying proteomics data enriched using clinical information. The pipeline allows users to relate established blood biomarkers with clinical parameters and proteomics data. The proposed pipeline entails three main phases: (i) feature selection, (ii) models training, and (iii) models ensembling. We report the experience of applying such a pipeline to prostate-related diseases. Models have been trained on several biological datasets. We report experimental results about two datasets that result from the integration of clinical and mass spectrometry-based data in the contexts of serum and urine analysis. The pipeline receives input data from blood analytes, tissue samples, proteomic analysis, and urine biomarkers. It then trains different models for feature selection, classification and voting. The presented pipeline has been applied on two datasets obtained in a 2 years research project which aimed to extract hidden information from mass spectrometry, serum, and urine samples from hundreds of patients. We report results on analyzing prostate datasets serum with 143 samples, including 79 PCa and 84 BPH patients, and an urine dataset with 121 samples, including 67 PCa and 54 BPH patients. As results pipeline allowed to identify interesting peptides in the two datasets, 6 for the first one and 2 for the second one. The best model for both serum (AUC=0.87, Accuracy=0.83, F1=0.81, Sensitivity=0.84, Specificity=0.81) and urine (AUC=0.88, Accuracy=0.83, F1=0.83, Sensitivity=0.85, Specificity=0.80) datasets showed good predictive performances. We made the pipeline code available on GitHub and we are confident that it will be successfully adopted in similar clinical setups.

Indexed as

Prostatic HyperplasiaProstatic NeoplasmsBiomarkersHumansMachine LearningMalePeptidesProstateProteomicsBiomarkersPeptidesBiological pipelineData enhancingMachine learningProstate cancer

Identifiers

PMID38584282
PMCPMC11000316
OpenAlexW4394575124

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.