Evidence map›Paper›PMID 39737563›Full record

ArticleBriefings in bioinformatics2024

Dual-stage optimizer for systematic overestimation adjustment applied to multi-objective genetic algorithms for biomarker selection.

Luca Cattelani, Vittorio Fortino

Abstract read
In one paragraph

Article in Briefings in bioinformatics, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Luca CattelaniSchool of Medicine, Institute of Biomedicine, University of Eastern Finland, Yliopistonranta 1, PO Box 1627, 70211 Kuopio, Finland.ORCID 0000-0003-4852-2310
Vittorio FortinoSchool of Medicine, Institute of Biomedicine, University of Eastern Finland, Yliopistonranta 1, PO Box 1627, 70211 Kuopio, Finland.ORCID 0000-0001-8693-5285

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

The selection of biomarker panels in omics data, challenged by numerous molecular features and limited samples, often requires the use of machine learning methods paired with wrapper feature selection techniques, like genetic algorithms. They test various feature sets-potential biomarker solutions-to fine-tune a machine learning model's performance for supervised tasks, such as classifying cancer subtypes. This optimization process is undertaken using validation sets to evaluate and identify the most effective feature combinations. Evaluations have performance estimation error, measurable as discrepancy between validation and test set performance, and when the selection involves many models the best ones are almost certainly overestimated. This issue is also relevant in a multi-objective feature selection process where various characteristics of the biomarker panels are optimized, such as predictive performances and feature set size. Methods have been proposed to reduce the overestimation after a model has already been selected in single-objective problems, but no algorithm existed capable of reducing the overestimation during the optimization, improving model selection, or applied in the more general multi-objective domain. We propose Dual-stage Optimizer for Systematic overestimation Adjustment in Multi-Objective problems (DOSA-MO), a novel multi-objective optimization wrapper algorithm that learns how the original estimation, its variance, and the feature set size of the solutions predict the overestimation. DOSA-MO adjusts the expectation of the performance during the optimization, improving the composition of the solution set. We verify that DOSA-MO improves the performance of a state-of-the-art genetic algorithm on left-out or external sample sets, when predicting cancer subtypes and/or patient overall survival, using three transcriptomics datasets for kidney and breast cancer.

Indexed as

AlgorithmsBiomarkers, TumorMachine LearningComputational BiologyHumansNeoplasmsBiomarkers, Tumorbiomarker discoverygenetic algorithmsmodel selectionmulti-objectiveomicsoverestimation adjustment

Identifiers

PMID39737563
PMCPMC11684899

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.