Evidence map›Paper›PMID 40894641›Full record

ArticlebioRxiv : the preprint server for biology2025

Pre-training Genomic Language Model with Variants for Better Modeling Functional Genomics.

Tianyu Liu, Xiangyu Zhang, Jiecong Lin, Luca Pinello, Rex Ying, Hongyu Zhao

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

6 authors.

Tianyu LiuInterdepartmental Program in Computational Biology & Bioinformatics, Yale University, New Haven, 06511, CT, USA.
Xiangyu ZhangDepartment of Biostatistics, Yale University, New Haven, 06511, CT, USA.
Jiecong LinMassachusetts General Hospital, Cambridge, 02129, MA, USA.
Luca PinelloBroad Institute of MIT and Harvard, Cambridge, 02142, MA, USA.ORCID 0000-0003-1109-3823
Rex YingDepartment of Computer Science, Yale University, New Haven, 06511, CT, USA.
Hongyu ZhaoInterdepartmental Program in Computational Biology & Bioinformatics, Yale University, New Haven, 06511, CT, USA.ORCID 0000-0003-1195-9607

Funding

Laboratory, Data Analysis, and Coordinating Center (LDACC) for the Developmental Human Genotype-Tissue Expression ProjectU24HG012108 · NHGRI · YALE UNIVERSITY · PI GERSTEIN, MARK BENDER, HUTTNER, ANITA JULIANE · 2021 to 2025
$8.7M
Computational and Statistical Methods to determine variant effect across cell types and development stagesU01HG013840 · NHGRI · YALE UNIVERSITY · PI GERSTEIN, MARK BENDER, ZHAO, HONGYU · 2024 to 2024
$1.9M
NHGRI NIH HHS U01 HG013840NHGRI NIH HHS U24 HG012108
6 · The paper itself

Abstract

Sequence-to-function models can predict gene expression from sequence data and be used to link genetic information with transcriptomics data to understand regulatory processes and their effects on complex phenotypes. The genomic language models are pre-trained with large-scale DNA sequences and can generate robust representations of these sequences by learning the genomic context. However, few studies can estimate the predictability of gene expression levels and bridge these two classes of models together to explore individualized gene expression prediction. In this manuscript, we propose UKBioBERT as a DNA language model pre-trained with genetic variants from UK BioBank. We demonstrate that UKBioBERT generates informative embeddings capable of identifying gene functions, and improving gene expression prediction in cell lines, thereby enhancing our understanding of gene expression predictability. Building upon these embeddings, we combine UKBioBERT with state-of-the-art sequence-to-function architectures, Enformer and Borzoi, to create UKBioFormer and UKBioZoi. These models exhibit better performance in predicting highly predictable gene expression levels and can be generalized across different cohorts. Furthermore, UKBioFormer effectively captures the relationship between genetic variants and expression variations, enabling in-silico mutation analyses and eQTL identification. Collectively, our findings underscore the value of integrating genomic language models and sequence-to-function approaches for advancing functional genomics research.

Indexed as

DNA Sequence ModelFunctional GenomicsGene ExpressionLarge Language ModelSequence-to-Function Models

Identifiers

PMID40894641
PMCPMC12393247

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.