Evidence map›Paper›PMID 35336739›Full record

ArticleBiology2022

Machine Learning-Based Identification of Colon Cancer Candidate Diagnostics Genes.

Saraswati Koppad, Annappa Basava, Katrina Nash, Georgios V Gkoutos, Animesh Acharjee

Abstract read
In one paragraph

Article in Biology, 2022. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 21 papers.

0numbers the graph read from it
0cells of the map it votes in
21citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

21 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Saraswati KoppadDepartment of Computer Science and Engineering, National Institute of Technology Karnataka, Mangalore 575025, India.ORCID 0000-0003-4547-0551
Annappa BasavaDepartment of Computer Science and Engineering, National Institute of Technology Karnataka, Mangalore 575025, India.
Katrina NashCollege of Medical and Dental Sciences, University of Birmingham, Birmingham B15 2TT, UK.ORCID 0000-0002-5204-9688
Georgios V GkoutosInstitute of Cancer and Genomic Sciences, University of Birmingham, Birmingham B15 2TT, UK.ORCID 0000-0002-2061-091X
Animesh AcharjeeInstitute of Cancer and Genomic Sciences, University of Birmingham, Birmingham B15 2TT, UK.ORCID 0000-0003-2735-7010

Funding

Medical Research Council HDRUK/CFC/01
6 · The paper itself

Abstract

backgroundColorectal cancer (CRC) is the third leading cause of cancer-related death and the fourth most commonly diagnosed cancer worldwide. Due to a lack of diagnostic biomarkers and understanding of the underlying molecular mechanisms, CRC's mortality rate continues to grow. CRC occurrence and progression are dynamic processes. The expression levels of specific molecules vary at various stages of CRC, rendering its early detection and diagnosis challenging and the need for identifying accurate and meaningful CRC biomarkers more pressing. The advances in high-throughput sequencing technologies have been used to explore novel gene expression, targeted treatments, and colon cancer pathogenesis. Such approaches are routinely being applied and result in large datasets whose analysis is increasingly becoming dependent on machine learning (ML) algorithms that have been demonstrated to be computationally efficient platforms for the identification of variables across such high-dimensional datasets.

methodsWe developed a novel ML-based experimental design to study CRC gene associations. Six different machine learning methods were employed as classifiers to identify genes that can be used as diagnostics for CRC using gene expression and clinical datasets. The accuracy, sensitivity, specificity, F1 score, and area under receiver operating characteristic (AUROC) curve were derived to explore the differentially expressed genes (DEGs) for CRC diagnosis. Gene ontology enrichment analyses of these DEGs were performed and predicted gene signatures were linked with miRNAs.

resultsWe evaluated six machine learning classification methods (Adaboost, ExtraTrees, logistic regression, naïve Bayes classifier, random forest, and XGBoost) across different combinations of training and test datasets over GEO datasets. The accuracy and the AUROC of each combination of training and test data with different algorithms were used as comparison metrics. Random forest (RF) models consistently performed better than other models. In total, 34 genes were identified and used for pathway and gene set enrichment analysis. Further mapping of the 34 genes with miRNA identified interesting miRNA hubs genes.

conclusionsWe identified 34 genes with high accuracy that can be used as a diagnostics panel for CRC.

Indexed as

biomarker identificationmachine learningpredictiontranscriptomicsvariable selection

Identifiers

PMID35336739
PMCPMC8944988

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.