ReviewAmerican journal of human genetics2025
A data model for population descriptors in genomic research.
Review in American journal of human genetics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
- Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors.Journal of medical Internet research · 2026Article
- The Continuity Trap in Data Science Health Research.Journal of medical Internet research · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
26 authors.
Funding
Abstract
Population descriptors used in genetic studies have broad social and translational implications. There are no globally agreed-upon definitions or usages of common population descriptors (e.g., race, ethnicity, nationality, and tribe), many of which are applied ad hoc and/or derived from political or bureaucratic conventions. Recent recommendations have encouraged the retention of as much granularity in population descriptors as possible during data preparation, analysis, and interpretation of research results. However, genomic research infrastructures (i.e., current practices, resources, and workflows in genomic research) often lack systematic and flexible organization, structure, and harmonization of multifaceted and detailed population descriptor data. This can lead to loss of information, barriers to international collaboration, and potential issues in clinical translation. Here, we describe a data model, developed by the NIH-funded Polygenic Risk Methods in Diverse Populations (PRIMED) Consortium, that organizes and retains detailed population descriptor data for future research use. The model supports a versatile, traceable, and reproducible harmonization system that offers multiple benefits over existing data structures. This data model affords researchers the flexibility to thoughtfully choose and scientifically justify their choice of population descriptors. It avoids the conflation of social identities with biological categories and guards against harmful typological inferences. Genomic research tools of this kind will be crucial for producing scientifically robust findings that minimize potential harms of descriptor misuse while maximizing benefits for diverse communities.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.