Evidence map›Paper›PMID 42581006›Full record

ArticleGenome research2026

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

Jonathan D Rosen, Arjun Devadas Vasanthakumari, Kilian Salomon, Nikola de Lange, Pyaree Mohan Dash, Pia Keukeleire, Ali Hassan, Alejandro Barrera, Beniamin Krupkin, Grace Oualline and 3 more

Abstract read
In one paragraph

Article in Genome research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Massively parallel reporter assays forbioRxiv : the preprint server for biology · 2026
    Article
  3. Promoter mutagenesis and a massively parallel reporter screen of thebioRxiv : the preprint server for biology · 2026
    Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

13 authors.

Jonathan D RosenDepartment of Genetics, University of North Carolina, Chapel Hill, North Carolina 27599-7264, USA.ORCID http://orcid.org/0000-0001-6396-4219
Arjun Devadas VasanthakumariInstitute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany.ORCID http://orcid.org/0000-0002-3090-1200
Kilian SalomonBerlin Institute of Health at Charité-Universitätsmedizin Berlin, 10117 Berlin, Germany.ORCID http://orcid.org/0009-0009-3182-7987
Nikola de LangeInstitute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany.ORCID http://orcid.org/0000-0002-8395-9369
Pyaree Mohan DashBerlin Institute of Health at Charité-Universitätsmedizin Berlin, 10117 Berlin, Germany.ORCID http://orcid.org/0000-0002-1005-0437
Pia KeukeleireInstitute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany.ORCID http://orcid.org/0000-0002-8828-636X
Ali HassanInstitute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany.ORCID http://orcid.org/0009-0000-2852-092X
Alejandro BarreraDepartment of Biostatistics and Bioinformatics, Duke University Medical School, Durham, North Carolina 27710, USA.
Beniamin KrupkinDepartment of Bioengineering and Therapeutic Sciences, University of California San Francisco, San Francisco, California 94143, USA.ORCID http://orcid.org/0000-0002-7457-7342
Grace OuallineDepartment of Genetics, Yale School of Medicine, New Haven, Connecticut 06510, USA.ORCID http://orcid.org/0009-0004-2486-7220
Martin KircherInstitute of Human Genetics, University Hospital Schleswig-Holstein, University of Luebeck, 23562 Lübeck, Germany.ORCID http://orcid.org/0000-0001-9278-5471
Michael I LoveDepartment of Genetics, University of North Carolina, Chapel Hill, North Carolina 27599-7264, USA.ORCID http://orcid.org/0000-0001-8401-0545
Max SchubachBerlin Institute of Health at Charité-Universitätsmedizin Berlin, 10117 Berlin, Germany; max.schubach@bih-charite.de.ORCID http://orcid.org/0000-0002-2032-6679

Funding

UNC-CH CENTER FOR ENVIRONMENTAL HEALTH &SUSCEPTIBILITYP30ES010126 · NIEHS · UNIV OF NORTH CAROLINA CHAPEL HILL · PI Hazel B Nichols · 2001 to 2026
$36.3M
Systematic in vivo characterization of disease-associated regulatory variantsUM1HG012003 · NHGRI · UNIV OF NORTH CAROLINA CHAPEL HILL · PI Michael Isaiah Love, KAREN L. MOHLKE · 2021 to 2026
$9.9M
Massively parallel characterization of variants and elements impacting transcriptional regulation in dynamic cellular systemsUM1HG011966 · NHGRI · UNIVERSITY OF WASHINGTON · PI Nadav Ahituv, Jay Ashok Shendure · 2021 to 2026
$9.7M
Multi-scale functional dissection and modeling of regulatory variation associated with human traitsR01HG012872 · NHGRI · YALE UNIVERSITY · PI Steven K. Reilly · 2023 to 2026
$3.1M
NHGRI NIH HHS R01 HG012872NHGRI NIH HHS UM1 HG011966NHGRI NIH HHS UM1 HG012003NIEHS NIH HHS P30 ES010126
6 · The paper itself

Abstract

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Indexed as

Genes, ReporterGenomicsHigh-Throughput Nucleotide SequencingSoftwareGenome, HumanHumansReproducibility of Results

Identifiers

PMID42581006
PMCPMC13534247

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.