Evidence map›Paper›PMID 30577846›Full record

ArticleBMC systems biology2018

Identification of gene signatures from RNA-seq data using Pareto-optimal cluster algorithm.

Saurav Mallik, Zhongming Zhao

Abstract read
In one paragraph

Article in BMC systems biology, 2018. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Saurav MallikCenter for Precision Health, School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, 77030, TX, USA.
Zhongming ZhaoCenter for Precision Health, School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, 77030, TX, USA. Zhongming.Zhao@uth.tmc.edu.

Funding

Transforming dbGaP genetic and genomic data to FAIR-ready by artificial intelligence and machine learning algorithmsR01LM012806 · NLM · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI Zhongming Zhao · 2017 to 2026
$3.7M
Transcriptional and Post-transcriptional Co-regulation in Lip DevelopmentR03DE027393 · NIDCR · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI IWATA, JUNICHI, ZHAO, ZHONGMING · 2018 to 2019
$308k
Mining Genomic Data in FaceBase for Cleft GenesR03DE028103 · NIDCR · UNIVERSITY OF TEXAS HLTH SCI CTR HOUSTON · PI ZHAO, ZHONGMING · 2018 to 2019
$308k
NIDCR NIH HHS R03 DE027393NIDCR NIH HHS R03 DE028103NLM NIH HHS R01 LM012806
6 · The paper itself

Abstract

backgroundGene signatures are important to represent the molecular changes in the disease genomes or the cells in specific conditions, and have been often used to separate samples into different groups for better research or clinical treatment. While many methods and applications have been available in literature, there still lack powerful ones that can take account of the complex data and detect the most informative signatures.

methodsIn this article, we present a new framework for identifying gene signatures using Pareto-optimal cluster size identification for RNA-seq data. We first performed pre-filtering steps and normalization, then utilized the empirical Bayes test in Limma package to identify the differentially expressed genes (DEGs). Next, we used a multi-objective optimization technique, "Multi-objective optimization for collecting cluster alternatives" (MOCCA in R package) on these DEGs to find Pareto-optimal cluster size, and then applied k-means clustering to the RNA-seq data based on the optimal cluster size. The best cluster was obtained through computing the average Spearman's Correlation Score among all the genes in pair-wise manner belonging to the module. The best cluster is treated as the signature for the respective disease or cellular condition.

resultsWe applied our framework to a cervical cancer RNA-seq dataset, which included 253 squamous cell carcinoma (SCC) samples and 22 adenocarcinoma (ADENO) samples. We identified a total of 582 DEGs by Limma analysis of SCC versus ADENO samples. Among them, 260 are up-regulated genes and 322 are down-regulated genes. Using MOCCA, we obtained seven Pareto-optimal clusters. The best cluster has a total of 35 DEGs consisting of all-upregulated genes. For validation, we ran PAMR (prediction analysis for microarrays) classifier on the selected best cluster, and assessed the classification performance. Our evaluation, measured by sensitivity, specificity, precision, and accuracy, showed high confidence.

conclusionsOur framework identified a multi-objective based cluster that is treated as a signature that can classify the disease and control group of samples with higher classification performance (accuracy 0.935) for the corresponding disease. Our method is useful to find signature for any RNA-seq or microarray data.

Indexed as

AlgorithmsSequence Analysis, RNACluster AnalysisComputational BiologyGene Expression ProfilingCervical cancerGene signatureK-meansLimmaPareto optimal clustering

Identifiers

PMID30577846
PMCPMC6302366

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.