Evidence map›Paper›PMID 34696650›Full record

ArticleStatistical methods in medical research2021

Variable selection with missing data in both covariates and outcomes: Imputation and machine learning.

Liangyuan Hu, Jung-Yi Joyce Lin, Jiayi Ji

Abstract read
In one paragraph

Article in Statistical methods in medical research, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.

0numbers the graph read from it
0cells of the map it votes in
11citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

11 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Using Tree-Based Machine Learning for Health Studies: Literature Review and Case Series.International journal of environmental research and public health · 2022
    Review
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Liangyuan HuDepartment of Biostatistics and Epidemiology, 242612Rutgers University School of Public Health, USA.ORCID 0000-0002-4067-892X
Jung-Yi Joyce LinDepartment of Population Health Science & Policy, 5925Icahn School of Medicine at Mount Sinai, USA.
Jiayi JiDepartment of Population Health Science & Policy, 5925Icahn School of Medicine at Mount Sinai, USA.

Funding

THE TISCH CANCER INSTITUTE - CANCER CENTER SUPPORT GRANTP30CA196521 · NCI · ICAHN SCHOOL OF MEDICINE AT MOUNT SINAI · PI Ramon E Parsons · 2015 to 2026
$35.4M
Flexible Bayesian approaches to causal inference with multilevel survival data and multiple treatmentsR21CA245855 · NCI · RBHS-SCHOOL OF PUBLIC HEALTH · PI HU, LIANGYUAN · 2020 to 2020
$459k
NCI NIH HHS P30 CA196521NCI NIH HHS R21 CA245855
6 · The paper itself

Abstract

Variable selection in the presence of both missing covariates and outcomes is an important statistical research topic. Parametric regression are susceptible to misspecification, and as a result are sub-optimal for variable selection. Flexible machine learning methods mitigate the reliance on the parametric assumptions, but do not provide as naturally defined variable importance measure as the covariate effect native to parametric models. We investigate a general variable selection approach when both the covariates and outcomes can be missing at random and have general missing data patterns. This approach exploits the flexibility of machine learning models and bootstrap imputation, which is amenable to nonparametric methods in which the covariate effects are not directly available. We conduct expansive simulations investigating the practical operating characteristics of the proposed variable selection approach, when combined with four tree-based machine learning methods, extreme gradient boosting, random forests, Bayesian additive regression trees, and conditional random forests, and two commonly used parametric methods, lasso and backward stepwise selection. Numeric results suggest that, extreme gradient boosting and Bayesian additive regression trees have the overall best variable selection performance with respect to the

Indexed as

Machine LearningBayes TheoremFemaleHumansRisk Factorsbootstrap imputation,variable selectionMissing at randomtree ensemblevariable importance

Identifiers

PMID34696650
PMCPMC11181487

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.