Evidence map›Paper›PMID 40972583›Full record

ArticleCell genomics2025

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Xihao Li, Andrew R Wood, Yuxin Yuan, Manrui Zhang, Yushu Huang, Gareth Hawkes, Robin N Beaumont, Michael N Weedon, Wenyuan Li, Xiaoyu Li and 2 more

Abstract read
In one paragraph

Article in Cell genomics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Article
  3. Mechanisms of breast cancer dormancy in bone metastasis.Clinical & experimental metastasis · 2026
    Review
  4. Article
  5. Article
  6. Article
  7. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

12 authors.

Xihao LiDepartment of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA; Department of Genetics, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA. Electronic address: xihaoli@unc.edu.
Andrew R WoodDepartment of Biomedical and Clinical Sciences, Faculty of Health and Life Sciences, University of Exeter, Exeter, UK. Electronic address: a.r.wood@exeter.ac.uk.
Yuxin YuanSchool of Mathematics and Statistics and KLAS, Northeast Normal University, Changchun, Jilin, China.
Manrui ZhangDepartment of Sociology, Tsinghua University, Beijing, China.
Yushu HuangDepartment of Big Data in Health Science, School of Public Health and Center of Clinical Big Data and Analytics of the Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, Zhejiang, China.
Gareth HawkesDepartment of Biomedical and Clinical Sciences, Faculty of Health and Life Sciences, University of Exeter, Exeter, UK.
Robin N BeaumontDepartment of Biomedical and Clinical Sciences, Faculty of Health and Life Sciences, University of Exeter, Exeter, UK.
Michael N WeedonDepartment of Biomedical and Clinical Sciences, Faculty of Health and Life Sciences, University of Exeter, Exeter, UK.
Wenyuan LiDepartment of Big Data in Health Science, School of Public Health and Center of Clinical Big Data and Analytics of the Second Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, Zhejiang, China.
Xiaoyu LiDepartment of Sociology, Tsinghua University, Beijing, China.
Xihong LinDepartment of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA; Department of Statistics, Harvard University, Cambridge, MA, USA. Electronic address: xlin@hsph.harvard.edu.
Zilin LiSchool of Mathematics and Statistics and KLAS, Northeast Normal University, Changchun, Jilin, China. Electronic address: lizl@nenu.edu.cn.

Funding

Translating Molecular and Clinical Data to Population Lung Cancer Risk AssessmentU19CA203654 · NCI · UNIVERSITY OF NEW MEXICO HEALTH SCIS CTR · PI Rayjean J. Hung · 2017 to 2026
$23.7M
Statistical Methods for Analysis of Massive Genetic and Genomic Data in Cancer ResearchR35CA197449 · NCI · HARVARD UNIVERSITY D/B/A HARVARD SCHOOL OF PUBLIC HEALTH · PI XIHONG LIN · 2015 to 2026
$10.9M
Powering whole genome sequence-based genetic discovery for common human diseases- Extended 2021-2022.U01HG009088 · NHGRI · HARVARD SCHOOL OF PUBLIC HEALTH · PI LIN, XIHONG, NEALE, BENJAMIN MICHAEL · 2016 to 2021
$5.1M
Predictive Modeling of the Functional and Phenotypic Impacts of Genetic VariantsU01HG012064 · NHGRI · UNIV OF MASSACHUSETTS MED SCH WORCESTER · PI Manuel Garber, XIHONG LIN · 2021 to 2026
$4.0M
Construction and Application of Comprehensive Knowledge Graphs for Alzheimer's DiseaseR01AG085581 · NIA · UNIV OF NORTH CAROLINA CHAPEL HILL · PI Yun Li, Hongtu Zhu · 2024 to 2026
$3.7M
Statistical Methods for Integrative Analysis of Large-Scale Whole Genome Sequencing Studies and Biobanks of Common DiseasesR01HL163560 · NHLBI · HARVARD UNIVERSITY D/B/A HARVARD SCHOOL OF PUBLIC HEALTH · PI XIHONG LIN · 2022 to 2026
$2.6M
Enhanced Machine Learning Tools for Complex Data Evaluation and Integration in Advancing Health OutcomesR01HL173044 · NHLBI · UNIV OF NORTH CAROLINA CHAPEL HILL · PI Baiming Zou, Fei Zou · 2025 to 2026
$1.3M
Medical Research Council MR/Y003748/1Medical Research Council UKRI327NCI NIH HHS R35 CA197449NCI NIH HHS U19 CA203654NHGRI NIH HHS U01 HG009088NHGRI NIH HHS U01 HG012064NHLBI NIH HHS R01 HL163560NHLBI NIH HHS R01 HL173044NIA NIH HHS R01 AG085581
6 · The paper itself

Abstract

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Indexed as

Biological Specimen BanksData ManagementGenome, HumanGenomicsWhole Genome SequencingDatabases, GeneticHumansSoftwareUK BiobankUnited Kingdomannotated genomic data structurebig data managementcloud computingfunctional annotationsfunctionally informed association analysesvcf2agds toolkitwhole-genome sequencing

Identifiers

PMID40972583
PMCPMC12802598

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.