Evidence map›Paper›PMID 41929314›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Integrating 730,947 exome sequences with clinical literature improves gene discovery.

Jeremy Guez, Julia K Goodrich, Mikhail A Moldovan, Katherine R Chao, Prathitha Kar, Ruchit Panchal, Michael W Wilson, Kristen M Laricchia, Greg Rohlicek, Dmitry Biba and 45 more

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

55 authors.

Jeremy GuezAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Julia K GoodrichCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Mikhail A MoldovanDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Katherine R ChaoProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Prathitha KarDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Ruchit PanchalCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Michael W WilsonAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Kristen M LaricchiaProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Greg RohlicekAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Dmitry BibaDepartment of Quantitative Biology, Cold Spring Harbor Laboratories, Cold Spring Harbor, NY, USA.
Daniel MartenProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Qin HeProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Philip W DarnowskyProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Riley GrantProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Ben WeisburdCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Samantha M BaxterProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Joshua NadeauProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Wenhan LuAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Steve JahlProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Sophie ParsaAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Abdallah LamaneHarvard-MIT Health Science Technology, Boston, MA, USA.
Stephanie DiTroiaProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Jack FuCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Xuefang ZhaoCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Elissa AlarmaniProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Charlotte TolonenBroad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Sam NovodBroad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Sam BryantProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Christine StevensProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Sinéad B ChapmanProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Caroline CusickProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Christopher VittalProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Laura D GauthierBroad Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Jacqueline I GoldsteinProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Daniel GoldsteinProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Daniel KingProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Timothy PoterbaProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Grace TiaoProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
gnomAD Project Consortium
Matteo TrancheroDepartment of ManagementManagement, University of Pennsylvania, Philadelphia, PA, USA.
William LotterDepartment of Data Science, Dana-Farber Cancer Institute, Boston, MA, USA.
Daniel G MacArthurCentre for Population Genomics, Garvan Institute of Medical Research and UNSW Sydney, Sydney, New South Wales, Australia.
Harrison BrandCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Vladimir SeplyarskiyDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Evan KochDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Michael E TalkowskiCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Matthew SolomonsonProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
Benjamin M NealeAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Anne O'Donnell-LuriaCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Hilary K FinucaneAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Shamil R SunyaevDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA, USA.
Mark J DalyAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.
Heidi L RehmCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Kaitlin E SamochaCenter for Genomic Medicine, Massachusetts General Hospital, Boston, MA, USA.
Konrad J KarczewskiAnalytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA, USA.

Funding

The Genome Aggregation Database (gnomAD)U24HG011450 · NHGRI · BROAD INSTITUTE, INC. · PI Mark Joseph Daly, Konrad Karczewski · 2021 to 2026
$22.3M
Broad Institute Mendelian Genomic Research CenterU01HG011755 · NHGRI · BROAD INSTITUTE, INC. · PI Anne O'Donnell-Luria, MICHAEL E TALKOWSKI · 2021 to 2026
$14.6M
Computational resources for genomic interpretation of type 2 diabetesU54DK105566 · NIDDK · BROAD INSTITUTE, INC. · PI MACARTHUR, DANIEL G, NEALE, BENJAMIN MICHAEL · 2014 to 2017
$8.6M
PRISM: Ethically-guided multimodal AI models for predicting disease pathogenesis in individuals with pathogenic variantsUG3HG014379 · NHGRI · MASSACHUSETTS GENERAL HOSPITAL · PI KARCZEWSKI, KONRAD, TATONETTI, NICHOLAS P · 2025 to 2025
$3.2M
Integrating genomic data and protein structures to improve measures of selective constraintR01HG012867 · NHGRI · MASSACHUSETTS GENERAL HOSPITAL · PI Kaitlin Elisabeth Samocha · 2024 to 2026
$2.5M
Integrating frequency and phenotypic association data to improve interpretation of rare genetic variantsR35GM157035 · NIGMS · MASSACHUSETTS GENERAL HOSPITAL · PI Konrad Karczewski · 2025 to 2026
$825k
NHGRI NIH HHS R01 HG012867NHGRI NIH HHS U01 HG011755NHGRI NIH HHS U24 HG011450NHGRI NIH HHS UG3 HG014379NIDDK NIH HHS U54 DK105566NIGMS NIH HHS R35 GM157035
6 · The paper itself

Abstract

Accurate estimates of allele frequencies aid in genetic discovery, including rare disease diagnosis, common disease investigations, and population genetics. Here, we present the Genome Aggregation Database version 4 (gnomAD v4), including 730,947 with exome sequences, a fivefold increase over previous releases. We demonstrate that statistical power to detect strong selective constraint continues to increase with sample size. We develop a new loss-of-function annotation pipeline, which learns genomic features predictive of nonsense-mediated decay and splicing effects from selection signals, achieving 90% precision for distinguishing likely true versus false positive loss-of-function variants. This improved pipeline, along with incorporation of highly deleterious missense variants into measures of loss-of-function intolerance, improves disease gene detection, particularly for short genes and those with gain-of-function mechanisms. To improve disease gene prediction, we systematically extract gene-disease associations from biomedical literature, map these to gene-level biological features, and integrate both with refined constraint metrics within a Bayesian framework, yielding state-of-the-art prediction of gene-disease relevance. We highlight genes under strong constraint but with limited clinical characterization, which are enriched in embryonic lethal and fertility phenotypes, thus prioritizing previously under-characterized disease genes. Together, these advances establish a unified framework for accelerating gene discovery and improving rare disease diagnosis.

Identifiers

PMID41929314
PMCPMC13042128

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.