Evidence map›Paper›PMID 39574876›Full record

ArticlemedRxiv : the preprint server for health sciences2024

Reducing Information and Selection Bias in EHR-Linked Biobanks via Genetics-Informed Multiple Imputation and Sample Weighting.

Maxwell Salvatore, Ritoban Kundu, Jiacong Du, Christopher R Friese, Alison M Mondul, David Hanauer, Haidong Lu, Celeste Leigh Pearce, Bhramar Mukherjee

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Maxwell SalvatoreDepartment of Epidemiology, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0002-3659-1514
Ritoban KunduCenter for Precision Health Data Science, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0003-0967-6755
Jiacong DuCenter for Precision Health Data Science, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0002-7768-9083
Christopher R FrieseRogel Cancer Center, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0002-2281-2056
Alison M MondulDepartment of Epidemiology, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0002-8843-1416
David HanauerDepartment of Learning Health Sciences, University of Michigan Medical School, Ann Arbor, MI, USA.
Haidong LuSection of General Internal Medicine, Department of Internal Medicine, Yale School of Medicine, New Haven, CT, USA.
Celeste Leigh PearceDepartment of Epidemiology, University of Michigan, Ann Arbor, MI, USA.ORCID 0000-0002-8023-7846
Bhramar MukherjeeDepartment of Biostatistics, Yale University, New Haven, CT, USA.ORCID 0000-0003-0118-4561

Funding

XenograftP30CA046592 · NCI · UNIVERSITY OF MICHIGAN AT ANN ARBOR · PI Eric R. Fearon · 1988 to 2026
$178.2M
NCI NIH HHS P30 CA046592
6 · The paper itself

Abstract

Electronic health records (EHRs) are valuable for public health and clinical research but are prone to many sources of bias, including missing data and non-probability selection. Missing data in EHRs is complex due to potential non-recording, fragmentation, or clinically informative absences. This study explores whether polygenic risk score (PRS)-informed multiple imputation for missing traits, combined with sample weighting, can mitigate missing data and selection biases in estimating disease-exposure associations. Simulations were conducted for missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR) conditions under different sampling mechanisms. PRS-informed multiple imputation showed generally lower bias, particularly when combined with sample weighting. For example, in biased samples of 10,000 with exposure and outcome MAR data, PRS-informed imputation had lower percent bias (3.8%) and better coverage rate (0.883) compared to PRS-uninformed (4.5%; 0.877) and complete case analyses (10.3%; 0.784) in covariate-adjusted, weighted, multiple imputation scenarios. In a case study using Michigan Genomics Initiative (n=50,026) data, PRS-informed imputation aligned more closely with a sample-weighted All of Us-derived benchmark than analyses ignoring missing data and selection bias. Researchers should consider leveraging genetic data and sample weighting to address biases from missing data and non-probability sampling in biobanks.

Indexed as

biobankelectronic health recordsexposomemissing dataselection bias

Identifiers

PMID39574876
PMCPMC11581092

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.