Evidence map›Paper›PMID 41024427›Full record

ArticleBiostatistics (Oxford, England)2025

Semiparametric efficient estimation of small genetic effects in large-scale population cohorts.

Olivier Labayle, Breeshey Roskams-Hieter, Joshua Slaughter, Kelsey Tetley-Campbell, Mark J van der Laan, Chris P Ponting, Sjoerd V Beentjes, Ava Khamseh

Abstract read
In one paragraph

Article in Biostatistics (Oxford, England), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Epistatic contributions to human traits via transcription factor mechanisms.medRxiv : the preprint server for health sciences · 2025
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Olivier LabayleSchool of Informatics, University of Edinburgh, 10 Crichton Street, Edinburgh EH8 9AB, United Kingdom.ORCID 0000-0002-3708-3706
Breeshey Roskams-HieterHealth Data Research UK, 215 Euston Road, London NW1 2BE, United Kingdom.ORCID 0000-0002-1119-2576
Joshua SlaughterSchool of Informatics, University of Edinburgh, 10 Crichton Street, Edinburgh EH8 9AB, United Kingdom.
Kelsey Tetley-CampbellMRC Human Genetics Unit, Institute of Genetics and Cancer, University of Edinburgh, Crewe Road South, Edinburgh EH4 2XU, United Kingdom.
Mark J van der LaanSchool of Public Health, University of California, Berkeley, 2121 Berkeley Way, Berkeley, CA 94720, United States.
Chris P PontingMRC Human Genetics Unit, Institute of Genetics and Cancer, University of Edinburgh, Crewe Road South, Edinburgh EH4 2XU, United Kingdom.ORCID 0000-0003-0202-7816
Sjoerd V BeentjesMRC Human Genetics Unit, Institute of Genetics and Cancer, University of Edinburgh, Crewe Road South, Edinburgh EH4 2XU, United Kingdom.ORCID 0000-0002-7998-4262
Ava KhamsehSchool of Informatics, University of Edinburgh, 10 Crichton Street, Edinburgh EH8 9AB, United Kingdom.ORCID 0000-0001-5203-2205

Funding

Targeted Learning using adaptive designs for HIV Epidemic control in East AfricaR01AI074345 · NIAID · UNIVERSITY OF CALIFORNIA BERKELEY · PI PETERSEN, MAYA LIV, VANDERLAAN, MARK J · 2007 to 2023
$6.6M
Health Data Research UKInstitute of Genetics and CancerMedical Research Council MC_UU_00007/15Medical Research Council MC_UU_00009/2NIAID NIH HHS R01 AI074345NIH HHS R01AI074345UKRI Centre for Doctoral TrainingUnited Kingdom Research and Innovation EP/S02431X/1University of Edinburgh
6 · The paper itself

Abstract

Population genetics seeks to quantify DNA variant associations with traits or diseases, as well as interactions among variants and with environmental factors. Computing millions of estimates in large cohorts in which small effect sizes and tight confidence intervals are expected, necessitates minimizing model-misspecification bias to increase power and control false discoveries. We present TarGene, a unified statistical workflow for the semi-parametric efficient and double robust estimation of genetic effects including $ k $-point interactions among categorical variables in the presence of confounding and weak population dependence. $ k $-point interactions, or Average Interaction Effects (AIEs), are a direct generalization of the usual average treatment effect (ATE). We estimate genetic effects with cross-validated and/or weighted versions of Targeted Minimum Loss-based Estimators (TMLE) and One-Step Estimators (OSE). The effect of dependence among data units on variance estimates is corrected by using sieve plateau variance estimators based on genetic relatedness across the units. We present extensive realistic simulations to demonstrate power, coverage, and control of type I error. Our motivating application is the targeted estimation of genetic effects on trait, including two-point and higher-order gene-gene and gene-environment interactions, in large-scale genomic databases such as UK Biobank and All of Us. All cross-validated and/or weighted TMLE and OSE for the AIE $ k $-point interaction, as well as ATEs, conditional ATEs and functions thereof, are implemented in the general purpose Julia package TMLE.jl. For high-throughput applications in population genomics, we provide the open-source Nextflow pipeline and software TarGene which integrates seamlessly with modern high-performance and cloud computing platforms.

Indexed as

Genetics, PopulationModels, GeneticModels, StatisticalBiostatisticsCohort StudiesComputer SimulationHumansbioinformaticsnon-parametric methodsstatistical geneticsstatistical methods in epidemiology

Identifiers

PMID41024427
PMCPMC12479317

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.