Evidence map›Paper›PMID 41339306›Full record

ArticleNature communications2025

Long-read transcriptomics of a diverse human cohort reveals ancestry bias in gene annotation.

Pau Clavell-Revelles, Fairlie Reese, Sílvia Carbonell-Sala, Fabien Degalez, Carme Arnan, Winona Oliveros, Emilio Palumbo, Tamara Perteghella, Roderic Guigó, Marta Melé

Abstract read
In one paragraph

Article in Nature communications, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 10 papers.

0numbers the graph read from it
0cells of the map it votes in
10citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

10 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
  4. Article
  5. Review
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

10 authors.

Pau Clavell-Revelles *Department of Life Sciences, Barcelona Supercomputing Center (BCN-CNS), Barcelona, Catalonia, Spain.ORCID http://orcid.org/0009-0003-1269-293X
Fairlie Reese *Department of Life Sciences, Barcelona Supercomputing Center (BCN-CNS), Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0002-9240-0102
Sílvia Carbonell-SalaCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0001-7956-6215
Fabien DegalezCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0001-8252-6425
Carme ArnanCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0002-7431-2088
Winona OliverosDepartment of Life Sciences, Barcelona Supercomputing Center (BCN-CNS), Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0002-3646-8874
Emilio PalumboCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0003-4599-8161
Tamara PerteghellaCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain.ORCID http://orcid.org/0000-0001-7073-6361
Roderic GuigóCentre for Genomic Regulation (CRG), The Barcelona Institute of Science and Technology, Barcelona, Catalonia, Spain. roderic.guigo@crg.cat.ORCID http://orcid.org/0000-0002-5738-4477
Marta MeléDepartment of Life Sciences, Barcelona Supercomputing Center (BCN-CNS), Barcelona, Catalonia, Spain. marta.mele@bsc.es.ORCID http://orcid.org/0000-0001-8874-6453

Funding

GENCODE Resource ProjectU41HG007234 · NHGRI · SANGER INSTITUTE · PI FLICEK, PAUL · 2013 to 2020
$20.3M
Generating a full-length reference transcriptome for human protein-coding genesU24HG011451 · NHGRI · DANA-FARBER CANCER INST · PI David E. Hill, Marc Vidal · 2022 to 2026
$3.3M
Ministry of Economy and Competitiveness | Agencia Estatal de Investigación (Spanish Agencia Estatal de Investigación) RYC-2017-22249NHGRI NIH HHS U24 HG011451NHGRI NIH HHS U41 HG007234Wellcome Trust
6 · The paper itself

Abstract

Accurate gene annotations are fundamental for interpreting genetic variation, cellular function, and disease mechanisms. However, current human gene annotations are largely derived from transcriptomic data of individuals with European ancestry, leaving gaps of annotation that remain uncharacterized. Here, we generate over 800 million full-length reads with long-read RNA-seq in 43 lymphoblastoid cell line samples from eight genetically-diverse human populations and build a cross-ancestry gene annotation. We demonstrate that transcripts from non-European samples are underrepresented in reference gene annotations, leading to incomplete characterization in allele-specific transcript usage. Furthermore, we show that personal genome assemblies enhance transcript discovery compared to the generic GRCh38 reference assembly, even though genomic regions unique to each individual are heavily depleted of genes. These findings underscore the urgent need for a more inclusive gene annotation framework that accurately represents global transcriptome diversity.

Indexed as

Molecular Sequence AnnotationTranscriptomeAllelesCell LineCohort StudiesGene Expression ProfilingGenetic VariationGenome, HumanHumansRNA-Seq

Identifiers

PMID41339306
PMCPMC12675792

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.