Evidence map›Paper›PMID 41408184›Full record

ArticleBMC medical research methodology2025

A novel statistical feature selection framework for biomarker discovery and cancer classification via multiomics integration.

Moshira S Ghaleb, Maryam N Al-Berry, Hala M Ebied, Mohamed F Tolba

Abstract read
In one paragraph

Article in BMC medical research methodology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. AI based multiomics integration for cancer diagnosis and prognosis.Journal, genetic engineering & biotechnology · 2026
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Moshira S GhalebScientific Computing Department Faculty of Computer and Information Sciences, Ain Shams University, Cairo, Egypt. moshirasg@cis.asu.edu.eg.
Maryam N Al-BerryScientific Computing Department Faculty of Computer and Information Sciences, Ain Shams University, Cairo, Egypt.
Hala M EbiedScientific Computing Department Faculty of Computer and Information Sciences, Ain Shams University, Cairo, Egypt.
Mohamed F TolbaScientific Computing Department Faculty of Computer and Information Sciences, Ain Shams University, Cairo, Egypt.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundEarly cancer diagnosis is essential for improving prognosis and guiding treatment. However, the high dimensionality and complexity of omics data present major challenges. Computational approaches that extract stable biomarkers and enable reliable classification across cancer types and stages are needed.

methodsA novel feature selection method, sDCFE (synergistic Discriminative Cluster-based Feature Extraction), was developed by extending Fisher-like variance analysis with a median absolute deviation (MAD) regularization term and a cluster separation component to enhance robustness and interpretability. Features selected by sDCFE were compared with those obtained from XGBoost, and the intersected set of 82 genes was evaluated through functional enrichment (KEGG, Reactome, GO BP), survival analysis (Kaplan-Meier, Cox regression), and biomarker novelty assessment against six external resources. Hybrid classification models integrating XGBoost, sDCFE, and deep learning were applied to pancancer classification, and the framework was further extended to lung squamous cell carcinoma (LUSC) staging using RNA-seq and methylation data.

resultsThe overlap between sDCFE and XGBoost yielded 82 candidate biomarkers enriched in cancer-related pathways, including cell cycle regulation, immune signalling, and DNA repair. Novelty assessment stratified these genes into established, emerging, and novel categories. Six genes-HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP-emerged as the most promising candidates, supported by enrichment and survival associations across multiple cancers. The hybrid model achieved near-perfect pancancer classification on TCGA (accuracy = 99.3%, MCC = 0.992, AUC = 1.0) and demonstrated strong generalizability on PCAWG (accuracy = 94%, MCC = 0.929, AUC = 0.997). In the LUSC staging task, multiomics integration improved classification performance: the CNN-based model reached 84% accuracy, while logistic regression applied to sDCFE-ranked features achieved 88.5% accuracy with superior calibration, highlighting the robustness of the selected features.

conclusionsDCFE provides a principled extension of Fisher-like methods, enabling stable and interpretable biomarker selection. When combined with XGBoost and deep learning, the framework achieves highly accurate and biologically grounded cancer classification across both cancer types and stages. The identification of novel and prognostic biomarkers, including HFE2, LOC339674, SERINC2, SFTA3, SOX2OT, and ACPP, underscores its translational potential. These results position the framework as a promising precision oncology tool to support early diagnosis, risk stratification, and treatment decision-making.

Indexed as

Biomarkers, TumorComputational BiologyGenomicsLung NeoplasmsNeoplasmsAlgorithmsCluster AnalysisDeep LearningHumansMultiomicsPrognosisBiomarkers, TumorCancer stageLUSCMachine learningMethylationMultiomicsPancancerRNA-seqSDCFEStatisticalTCGA

Identifiers

PMID41408184
PMCPMC12822226

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.