Evidence map›Paper›PMID 41316728›Full record

ArticleNucleic acids research2026

PolyA_DB v4: systematic polyA site identification and isoform annotation in human and mouse genomes using 3' end and long-read sequencing data.

Shan Yu, Wei Chun Chen, Luyang Wang, San Jewell, Ayna Mammedova, Seong Woo Han, Jayamanna Wickramasinghe, Yoseph Barash, Bin Tian

Abstract read
In one paragraph

Article in Nucleic acids research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Shan YuGenome Regulation and Cell Signalling Program, Ellen and Ronald Caplan Cancer Center, The Wistar Institute, Philadelphia, PA 19104, United States.
Wei Chun ChenGenome Regulation and Cell Signalling Program, Ellen and Ronald Caplan Cancer Center, The Wistar Institute, Philadelphia, PA 19104, United States.
Luyang WangGenome Regulation and Cell Signalling Program, Ellen and Ronald Caplan Cancer Center, The Wistar Institute, Philadelphia, PA 19104, United States.
San JewellDepartment of Genetics, University of Pennsylvania, Philadelphia, PA 19104, United States.
Ayna MammedovaGenome Regulation and Cell Signalling Program, Ellen and Ronald Caplan Cancer Center, The Wistar Institute, Philadelphia, PA 19104, United States.
Seong Woo HanDepartment of Genetics, University of Pennsylvania, Philadelphia, PA 19104, United States.
Jayamanna WickramasingheCenter for Systems and Computational Biology, The Wistar Institute, Philadelphia, PA 19104, United States.
Yoseph BarashDepartment of Genetics, University of Pennsylvania, Philadelphia, PA 19104, United States.
Bin TianGenome Regulation and Cell Signalling Program, Ellen and Ronald Caplan Cancer Center, The Wistar Institute, Philadelphia, PA 19104, United States.ORCID 0000-0001-8903-8256

Funding

Tumor Microenvironment and MetastasisP30CA010815 · NCI · WISTAR INSTITUTE · PI Aaron Robert Goldman · 1985 to 2026
$75.9M
Regulation of Alternative Cleavage and PolyadenylationR01GM084089 · NIGMS · WISTAR INSTITUTE · PI TIAN, BIN · 2008 to 2023
$6.7M
Regulation of gene expression by alternative polyadenylationR35GM153277 · NIGMS · WISTAR INSTITUTE · PI BIN TIAN · 2024 to 2026
$2.0M
Identifying regulatory uORFs as a targetable axis for hereditary diseaseR01GM147739 · NIGMS · UNIVERSITY OF PENNSYLVANIA · PI BARASH, YOSEPH, HAND, NICHOLAS JOSEPH · 2022 to 2025
$1.7M
Methods for improving clinical diagnostic by detection, prediction, interpretation and prioritization of aberrant transcriptome variationsR01LM013437 · NLM · UNIVERSITY OF PENNSYLVANIA · PI BARASH, YOSEPH · 2020 to 2023
$1.4M
National Science Foundation Cooperative Agreement DBI-2400327NCI NIH HHS P30 CA010815NIGMS NIH HHS R01 GM084089NIGMS NIH HHS R01 GM147739NIGMS NIH HHS R35 GM153277NIH HHS R01GM084089NIH HHS R01GM147739NIH HHS R01LM013437NIH HHS R35GM153277NLM NIH HHS R01 LM013437
6 · The paper itself

Abstract

The cleavage and polyadenylation site (PAS) defines the 3' end of almost all protein-coding and long non-coding RNAs in eukaryotes. Most genes harbor multiple PAS, resulting in expression of alternative polyadenylation (APA) isoforms. Here, we present PolyA_DB version 4 (https://exon.apps.wistar.org/polya_db/v4/), an updated database dedicated to PAS in mammalian genomes. By exhaustive mining of human and mouse transcriptomic data sets generated by the 3' region extraction and deep sequencing plus (3'READS+) method, corresponding to ∼2.3 billion PAS-supporting reads for each species, we identify ∼1.4 million PAS in both human and mouse genomes, increasing PAS coverage over the last database version by 4.9- and 3.5-fold, respectively. Of the full PAS set (named Max collection), 20% of them match the transcript end sites (TES) of public long-read RNA sequencing (LR-RNA-seq) data. Notably, ∼10%-20% of LR-RNA-seq TES do not match our annotated PAS, suggesting 3' end artifacts derived plausibly from internal A-rich regions of RNA. However, LR-RNA-seq data substantially complement RefSeq-based assignment of PAS to genes and are highly valuable in subtyping APA events in the context of splicing configuration. PolyA_DB v4 also contains PAS conservation and PAS strength information and is linked to UCSC Genome Browser for data visualization.

Indexed as

Databases, GeneticGenome, HumanPoly APolyadenylationAnimalsGenomeHigh-Throughput Nucleotide SequencingHumansMiceMolecular Sequence AnnotationSequence Analysis, RNASoftwareTranscriptomePoly A

Identifiers

PMID41316728
PMCPMC12807684

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.