Evidence map›Paper›PMID 35313824›Full record

ArticleBMC bioinformatics2022

Evaluation of tree-based statistical learning methods for constructing genetic risk scores.

Michael Lau, Claudia Wigmann, Sara Kress, Tamara Schikowski, Holger Schwender

Abstract read
In one paragraph

Article in BMC bioinformatics, 2022. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Review
  5. Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Michael LauMathematical Institute, Heinrich Heine University, Düsseldorf, Germany. michael.lau@hhu.de.ORCID https://orcid.org/0000-0002-5327-8351
Claudia WigmannIUF - Leibniz Research Institute for Environmental Medicine, Düsseldorf, Germany.
Sara KressIUF - Leibniz Research Institute for Environmental Medicine, Düsseldorf, Germany.
Tamara SchikowskiIUF - Leibniz Research Institute for Environmental Medicine, Düsseldorf, Germany.
Holger SchwenderMathematical Institute, Heinrich Heine University, Düsseldorf, Germany.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundGenetic risk scores (GRS) summarize genetic features such as single nucleotide polymorphisms (SNPs) in a single statistic with respect to a given trait. So far, GRS are typically built using generalized linear models or regularized extensions. However, these linear methods are usually not able to incorporate gene-gene interactions or non-linear SNP-response relationships. Tree-based statistical learning methods such as random forests and logic regression may be an alternative to such regularized-regression-based methods and are investigated in this article. Moreover, we consider modifications of random forests and logic regression for the construction of GRS.

resultsIn an extensive simulation study and an application to a real data set from a German cohort study, we show that both tree-based approaches can outperform elastic net when constructing GRS for binary traits. Especially a modification of logic regression called logic bagging could induce comparatively high predictive power as measured by the area under the curve and the statistical power. Even when considering no epistatic interaction effects but only marginal genetic effects, the regularized regression method lead in most cases to inferior results.

conclusionsWhen constructing GRS, we recommend taking random forests and logic bagging into account, in particular, if it can be assumed that possibly unknown epistasis between SNPs is present. To develop the best possible prediction models, extensive joint hyperparameter optimizations should be conducted.

Indexed as

AlgorithmsPolymorphism, Single NucleotideCohort StudiesHumansRegression AnalysisRisk FactorsBaggingElastic netEpistasisLogic regressionPolygenic risk scoresRandom forestsSimulation studyStatistical learningVariable selection

Identifiers

PMID35313824
PMCPMC8935722

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.