ArticleBMC systems biology2018
Identification of gene signatures from RNA-seq data using Pareto-optimal cluster algorithm.
Article in BMC systems biology, 2018. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
10 citing papers in PubMed.
- Sample size requirements for machine learning classification of binary outcomes in bulk RNA-Seq data.BMC bioinformatics · 2026Article
- Investigating the overlap of machine learning algorithms in the final results of RNA-seq analysis on gene expression estimation.Health information science and systems · 2024Article
- Emerging Prognostic Markers in Patients Undergoing Liver Resection for Hepatocellular Carcinoma: A Narrative Review.Cancers · 2024Review
- Comparison of five supervised feature selection algorithms leading to top features and gene signatures from multi-omics data in cancer.BMC bioinformatics · 2022Article
- In silico ranking of phenolics for therapeutic effectiveness on cancer stem cells.BMC bioinformatics · 2020Article
- Identification of specific microRNA-messenger RNA regulation pairs in four subtypes of breast cancer.IET systems biology · 2020Article
- A Comparative Analysis of Single-Cell Transcriptome Identifies Reprogramming Driver Factors for Efficiency Improvement.Molecular therapy. Nucleic acids · 2020Article
- Aberrantly Methylated-Differentially Expressed Genes Identify Novel Atherosclerosis Risk Subtypes.Frontiers in genetics · 2020Article
- Multi-Objective Optimized Fuzzy Clustering for Detecting Cell Clusters from Single-Cell Expression Profiles.Genes · 2019Article
- The International Conference on Intelligent Biology and Medicine (ICIBM) 2018: systems biology on diverse data types.BMC systems biology · 2018Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
Abstract
backgroundGene signatures are important to represent the molecular changes in the disease genomes or the cells in specific conditions, and have been often used to separate samples into different groups for better research or clinical treatment. While many methods and applications have been available in literature, there still lack powerful ones that can take account of the complex data and detect the most informative signatures.
methodsIn this article, we present a new framework for identifying gene signatures using Pareto-optimal cluster size identification for RNA-seq data. We first performed pre-filtering steps and normalization, then utilized the empirical Bayes test in Limma package to identify the differentially expressed genes (DEGs). Next, we used a multi-objective optimization technique, "Multi-objective optimization for collecting cluster alternatives" (MOCCA in R package) on these DEGs to find Pareto-optimal cluster size, and then applied k-means clustering to the RNA-seq data based on the optimal cluster size. The best cluster was obtained through computing the average Spearman's Correlation Score among all the genes in pair-wise manner belonging to the module. The best cluster is treated as the signature for the respective disease or cellular condition.
resultsWe applied our framework to a cervical cancer RNA-seq dataset, which included 253 squamous cell carcinoma (SCC) samples and 22 adenocarcinoma (ADENO) samples. We identified a total of 582 DEGs by Limma analysis of SCC versus ADENO samples. Among them, 260 are up-regulated genes and 322 are down-regulated genes. Using MOCCA, we obtained seven Pareto-optimal clusters. The best cluster has a total of 35 DEGs consisting of all-upregulated genes. For validation, we ran PAMR (prediction analysis for microarrays) classifier on the selected best cluster, and assessed the classification performance. Our evaluation, measured by sensitivity, specificity, precision, and accuracy, showed high confidence.
conclusionsOur framework identified a multi-objective based cluster that is treated as a signature that can classify the disease and control group of samples with higher classification performance (accuracy 0.935) for the corresponding disease. Our method is useful to find signature for any RNA-seq or microarray data.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.