ArticleStatistical methods in medical research2021
Variable selection with missing data in both covariates and outcomes: Imputation and machine learning.
Article in Statistical methods in medical research, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 11 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
11 citing papers in PubMed.
- Multi-Level Variable Selection Using a BART-Enhanced Mixed-Effects Framework.Statistics in medicine · 2026Article
- Longitudinal modeling of Post-COVID-19 condition over three years: A machine learning approach using clinical, neuropsychological, and fluid markers.Scientific reports · 2026Article
- Examining the Trajectory of Health-Related Quality of Life among Coronavirus Disease Patients.Journal of general internal medicine · 2024Article
- A new method for clustered survival data: Estimation of treatment effect heterogeneity and variable selection.Biometrical journal. Biometrische Zeitschrift · 2024Article
- Using Tree-Based Machine Learning for Health Studies: Literature Review and Case Series.International journal of environmental research and public health · 2022Review
- A Flexible Approach for Assessing Heterogeneity of Causal Treatment Effects on Patient Survival Using Large Datasets with Clustered Observations.International journal of environmental research and public health · 2022Article
- A flexible approach for causal inference with multiple treatments and clustered survival outcomes.Statistics in medicine · 2022Article
- CIMTx: An R Package for Causal Inference with Multiple Treatments using Observational Data.The R journal · 2022Article
- A FLEXIBLE SENSITIVITY ANALYSIS APPROACH FOR UNMEASURED CONFOUNDING WITH MULTIPLE TREATMENTS AND A BINARY OUTCOME WITH APPLICATION TO SEER-MEDICARE LUNG CANCER DATA.The annals of applied statistics · 2022Article
- A flexible approach for variable selection in large-scale healthcare database studies with missing covariate and outcome data.BMC medical research methodology · 2022Article
- Estimating heterogeneous survival treatment effects of lung cancer screening approaches: A causal machine learning analysis.Annals of epidemiology · 2021Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
Abstract
Variable selection in the presence of both missing covariates and outcomes is an important statistical research topic. Parametric regression are susceptible to misspecification, and as a result are sub-optimal for variable selection. Flexible machine learning methods mitigate the reliance on the parametric assumptions, but do not provide as naturally defined variable importance measure as the covariate effect native to parametric models. We investigate a general variable selection approach when both the covariates and outcomes can be missing at random and have general missing data patterns. This approach exploits the flexibility of machine learning models and bootstrap imputation, which is amenable to nonparametric methods in which the covariate effects are not directly available. We conduct expansive simulations investigating the practical operating characteristics of the proposed variable selection approach, when combined with four tree-based machine learning methods, extreme gradient boosting, random forests, Bayesian additive regression trees, and conditional random forests, and two commonly used parametric methods, lasso and backward stepwise selection. Numeric results suggest that, extreme gradient boosting and Bayesian additive regression trees have the overall best variable selection performance with respect to the
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.