ArticleCell genomics2025
Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.
Article in Cell genomics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
7 citing papers in PubMed.
- Article
- MetaSTAARlite: an all-in-one tool for biobank-scale whole-genome sequencing meta-analysis.Nature computational science · 2026Article
- Mechanisms of breast cancer dormancy in bone metastasis.Clinical & experimental metastasis · 2026Review
- Joint modeling of whole-genome sequencing data for human height via approximate message passing.Cell genomics · 2026Article
- Integrating common and rare variants improves polygenic risk prediction across diverse populations.Nature communications · 2026Article
- Rare coding and noncoding variants map 1,342 diseases and biomarkers in 490,549 whole genomes.medRxiv : the preprint server for health sciences · 2026Article
- Health risks and genetic architecture of objectively measured multidimensional sleep health.Nature communications · 2025Article
Corrections and comments
- Update of
Authors and funding
12 authors.
Funding
Abstract
Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.