Evidence map›Paper›PMID 42723633›Full record

ArticleBioinformatics (Oxford, England)2026

Genomic language model for predicting enhancers and their allele-specific activity in the human genome.

Rekha Sathian, Pratik Dutta, Ferhat Ay, Ramana V Davuluri

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

4 authors.

Rekha SathianDepartment of Biomedical Informatics, Stony Brook University, Stony Brook, NY 11794, United States.
Pratik DuttaDepartment of Biomedical Informatics, Stony Brook University, Stony Brook, NY 11794, United States.
Ferhat AyCenters for Cancer Immunotherapy and Autoimmunity, La Jolla Institute for Immunology, La Jolla, CA 92037, United States.ORCID 0000-0002-0708-6914
Ramana V DavuluriDepartment of Biomedical Informatics, Stony Brook University, Stony Brook, NY 11794, United States.ORCID 0000-0002-7053-1064

Funding

Studying the function of human genetic variation in the light of 3D genome organizationR35GM128938 · NIGMS · LA JOLLA INSTITUTE FOR IMMUNOLOGY · PI Ferhat Ay · 2018 to 2026
$4.6M
National Library of Medicine/National Institutes of Health funding R01LM01372201National Library of Medicine/National Institutes of Health funding R35GM128938NIGMS NIH HHS R35 GM128938
6 · The paper itself

Abstract

motivationPredicting and deciphering the regulatory logic of enhancers remains a significant challenge due to their complex sequence features and the absence of consistent genetic or epigenetic signatures that distinguish them from other genomic regions. Existing machine learning methods capture nucleotide composition but often fail to model sequence context effectively.

resultsWe present DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. Using ENCODE registry of candidate cis-regulatory elements (cCREs), we curated a benchmark dataset, consisting of 21 926 enhancers of 201 bp length and 46 159 enhancers of 350 bp length, as positive instances. The best fine-tuned model achieved 88.05% accuracy and a Matthews correlation coefficient of 76% on an independent dataset. Genome-wide application identified 1 684 595 enhancer regions covering 26.65% of the human genome. By performing integrative analyses with DNABERT-based transcription factor models, we identify 2681 statistically significant loss-of-function and 1917 gain-of-function enhancer variants, which respectively alter the function of 1623 and 1247 ENCODE-cCRE enhancers. Similarly, we identify 4057 candidate de novo enhancers, created by 5464 gain-of-function variants. These genome-wide enhancer annotations and candidate genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies. AVAILABILITY AND IMPLEMENTATION: DNABERT-Enhancer is freely available at https://github.com/DavuluriLab/DNABERT-Enhancer; Trained model predictions can be explored interactively via the web application at https://dnabert-enhancer-datarepo.streamlit.app/. The fine-tuned models are archived and citable through Zenodo (https://doi.org/10.5281/zenodo.19157566).

Indexed as

AllelesEnhancer Elements, GeneticGenome, HumanGenomicsModels, GeneticHumans

Identifiers

PMID42723633
PMCPMC13579158

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.