Evidence map›Paper›PMID 42585225›Full record

ArticlePLoS computational biology2026

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

Erik D VonKaenel, Lisa M Bramer, Javier E Flores, Thomas O Metz, Ernesto S Nakayasu, Bobbie-Jo M Webb-Robertson

Abstract read
In one paragraph

Article in PLoS computational biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Erik D VonKaenelBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.
Lisa M BramerBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.ORCID 0000-0002-8384-1926
Javier E FloresBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.
Thomas O MetzBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.
Ernesto S NakayasuBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.
Bobbie-Jo M Webb-RobertsonBiological Science Division, Pacific Northwest National Laboratory, Richland, Washington, United States of America.ORCID 0000-0002-4744-2397

Funding

Limited Competition: Continued Follow-up of Subjects and Initiation of a Second Case-control Cohort in The Environmental Determinants of Diabetes in The Young Study (TEDDY)U01DK128847 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI JEFFREY P KRISCHER · 2021 to 2026
$79.2M
Data Coordinating Center (DCC)for the Consortium for Identification of EnvironmenUC4DK095300 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2011 to 2011
$56.0M
NIH Prior Approval Process ProfessionalUL1TR002535 · NCATS · UNIVERSITY OF COLORADO DENVER · PI SOKOL, RONALD J. · 2018 to 2022
$51.1M
Data Coordinating CenterU01DK063790 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2002 to 2006
$50.1M
Follow-up on Subjects, Integrative Data Analysis and Measurement of Viral Antibodies in The Environmental Determinants of Diabetes in The Young Study (TEDDY)U01DK124166 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2020 to 2022
$31.1M
NIDDK Follow-up on Subjects and Immunological Assessments in The Environmental Determinants of Diabetes In The Young Study (TEDDY) (UC4)UC4DK117483 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2017 to 2020
$30.4M
Establish a Data Coordinating Center (DCC) for the Consortium for Identification UC4DK100238 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2014 to 2014
$28.9M
The Study of Epigenetics and Infections in The Environmental Determinants of Diabetes in the Young (TEDDY)UC4DK112243 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2016 to 2016
$27.0M
Data Coordinating Center (DCC) for the Consortium for Identification of Environmental Determinants of Diabetes in the Young (TEDDY) StudyUC4DK106955 · NIDDK · UNIVERSITY OF SOUTH FLORIDA · PI KRISCHER, JEFFREY P · 2015 to 2015
$20.0M
The Environmental Determinants of Diabetes in Young (TEDDY0U01DK063863 · NIDDK · HOSPITAL DISTRICT OF SOUTHWEST FINLAND · PI TOPPARI, JORMA · 2003 to 2022
$13.5M
The Environmental Triggers of Diabetes (TEDDY) in SwedenU01DK063861 · NIDDK · UNIVERSITY OF WASHINGTON · PI LERNMARK, AKE · 2003 to 2022
$13.4M
THE TEDDY STUDY - COLORADO CLINICAL CENTERU01DK063821 · NIDDK · UNIVERSITY OF COLORADO DENVER · PI REWERS, MARIAN J · 2003 to 2022
$12.3M
NCATS NIH HHS UL1 TR000064NCATS NIH HHS UL1 TR002535NIDDK NIH HHS HHSN267200700014CNIDDK NIH HHS R01 DK138355NIDDK NIH HHS U01 DK063790NIDDK NIH HHS U01 DK063821NIDDK NIH HHS U01 DK063829NIDDK NIH HHS U01 DK063836NIDDK NIH HHS U01 DK063861NIDDK NIH HHS U01 DK063863NIDDK NIH HHS U01 DK063865NIDDK NIH HHS U01 DK124166NIDDK NIH HHS U01 DK127786NIDDK NIH HHS U01 DK128847NIDDK NIH HHS UC4 DK063821NIDDK NIH HHS UC4 DK063829NIDDK NIH HHS UC4 DK063836NIDDK NIH HHS UC4 DK063861NIDDK NIH HHS UC4 DK063863NIDDK NIH HHS UC4 DK063865NIDDK NIH HHS UC4 DK095300NIDDK NIH HHS UC4 DK100238NIDDK NIH HHS UC4 DK106955NIDDK NIH HHS UC4 DK112243NIDDK NIH HHS UC4 DK117483
6 · The paper itself

Abstract

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Indexed as

Diabetes Mellitus, Type 1AlgorithmsBenchmarkingClassification AlgorithmsComputational BiologyGenomicsHumansMachine LearningMultiomicsSoftware

Identifiers

PMID42585225
PMCPMC13465820

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.