Evidence map›Paper›PMID 42172141›Full record

ArticleDatabase : the journal of biological databases and curation2026

Large-scale manual curation and harmonization of metadata from metagenomic and cancer genomic repositories: challenges and solutions.

Kaelyn Long, Kai Gravel-Pucillo, Levi Waldron, Sean Davis, Sehyun Oh

Abstract read
In one paragraph

Article in Database : the journal of biological databases and curation, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. bioRxiv : the preprint server for biology · 2026
    Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

5 authors.

Kaelyn LongInstitute for Implementation Science in Population Health, City University of New York School of Public Health, New York, NY, United States.
Kai Gravel-PucilloInstitute for Implementation Science in Population Health, City University of New York School of Public Health, New York, NY, United States.
Levi WaldronInstitute for Implementation Science in Population Health, City University of New York School of Public Health, New York, NY, United States.ORCID 0000-0003-2725-0694
Sean DavisDepartments of Biomedical Informatics and Medicine, University of Colorado Anschutz School of Medicine, Denver, CO, United States.
Sehyun OhInstitute for Implementation Science in Population Health, City University of New York School of Public Health, New York, NY, United States.ORCID 0000-0002-9490-3061

Funding

Cancer Genomics:Integrative and Scalable Solutions in R / BioconductorU24CA180996 · NCI · ROSWELL PARK CANCER INSTITUTE CORP · PI MORGAN, MARTIN T, WALDRON, LEVI · 2014 to 2023
$6.9M
Exploiting Public Metagenomic Data to Uncover Cancer-Microbiome RelationshipsR01CA230551 · NCI · GRADUATE SCHOOL OF PUBLIC HEALTH AND HEALTH POLICY · PI Levi Waldron · 2020 to 2026
$3.9M
Cancer Genomics: Integrative and Scalable Solutions in R/BioconductorU24CA289073 · NCI · GRADUATE SCHOOL OF PUBLIC HEALTH AND HEALTH POLICY · PI Sean Davis, Levi Waldron · 2024 to 2026
$3.2M
NCI NIH HHS R01 CA230551NCI NIH HHS U24 CA180996NCI NIH HHS U24 CA289073NIH HHS 3U24CA180996-10S1NIH HHS U24CA289073
6 · The paper itself

Abstract

Public omics repositories contain vast amounts of valuable data, but their metadata suffers from extreme heterogeneity, unstandardized terminologies, and quality issues that severely limit data reusability and cross-study integration. While prospective metadata standards exist, the majority of published omics data remain in non-standardized formats requiring retrospective harmonization. We performed comprehensive manual curation and harmonization of metadata, such as participant characteristics and study conditions, from 212 027 omics samples across 468 studies in two repositories: curatedMetagenomicData (93 studies, 22 588 samples) and cBioPortal (375 studies, 189 438 samples). Through systematic ontology mapping, we consolidated redundant, dispersed information into far fewer harmonized columns, reduced unique values, and increased the completeness of major attributes. This curation process revealed common metadata quality issues, including typos, inconsistent terminologies, misplaced values, conflicting annotations, and inappropriately merged information across attributes. We document the challenges, decisions, and solutions during this large-scale metadata harmonization. The harmonized metadata, accessible through the OmicsMLRepoR Bioconductor package, enables repository-wide queries and cross-study analyses previously challenging with heterogeneous metadata. Our experience provides practical guidance for similar curation efforts and demonstrates the value of investing in retrospective metadata improvement for existing public omics resources.

Indexed as

Databases, GeneticData CurationGenomicsMetadataMetagenomicsNeoplasmsBiocurationHumans

Identifiers

PMID42172141
PMCPMC13196698

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.