Evidence map›Paper›PMID 41537311›Full record

ReviewBriefings in bioinformatics2026

A comprehensive survey of genome language models in bioinformatics.

Liyuan Shu, Jiao Tang, Xiaoyu Guan, Daoqiang Zhang

Abstract readReview
In one paragraph

Review in Briefings in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Review
  5. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Liyuan ShuCollege of Artificial Intelligence, Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing University of Aeronautics and Astronautics, No. 29 Jiangjun Avenue, Jiangning District, Nanjing, Jiangsu Province 211106, China.ORCID 0009-0005-6119-4368
Jiao TangCollege of Artificial Intelligence, Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing University of Aeronautics and Astronautics, No. 29 Jiangjun Avenue, Jiangning District, Nanjing, Jiangsu Province 211106, China.
Xiaoyu GuanCollege of Artificial Intelligence, Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing University of Aeronautics and Astronautics, No. 29 Jiangjun Avenue, Jiangning District, Nanjing, Jiangsu Province 211106, China.ORCID 0000-0003-3985-3900
Daoqiang ZhangCollege of Artificial Intelligence, Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education, Nanjing University of Aeronautics and Astronautics, No. 29 Jiangjun Avenue, Jiangning District, Nanjing, Jiangsu Province 211106, China.

Funding

Key Research and Development Plan of Jiangsu Province BE2022842National Key R&D Program of China 2023YFF1204803National Natural Science Foundation of China 62136004National Natural Science Foundation of China 62276130
6 · The paper itself

Abstract

Large language models have revolutionized natural language processing by effectively modeling complex semantics and capturing long-range contextual relationships. Inspired by these advancements, genome language models (gLMs) have recently emerged, conceptualizing DNA and RNA sequences as biological texts and enabling the identification of intricate genomic grammar and distant regulatory interactions. This review examines the need for gLMs, emphasizing their capacity to overcome the limitations of traditional deep learning approaches in genomic sequence characterization. We comprehensively survey contemporary gLM architectures, including Transformer models, Hyena convolutions, and state space models, as well as various sequence tokenization strategies, assessing their applicability, and effectiveness across diverse genomic applications. Additionally, we discuss foundational pretraining strategies and provide an overview of genomic pretraining datasets spanning multiple species and functional domains. We critically analyze evaluation methodologies, including supervised, zero-shot, and few-shot learning paradigms, as well as fine-tuning approaches. An extensive taxonomy of downstream tasks is presented, alongside a summary of existing benchmarks and emerging trends. Finally, we contemplate key challenges such as data scarcity, interpretability, and the computational demands of genomic modeling, and propose a roadmap to guide future advances in genome language modeling.

Indexed as

Computational BiologyGenomeGenomicsNatural Language ProcessingAnimalsDeep LearningHumansLarge Language Modelsfoundation modelsgenome language modelssequence tokenizationstate space modelstransformerszero-shot learning

Identifiers

PMID41537311
PMCPMC12805252

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.