Evidence map›Paper›PMID 41985060›Full record

ArticleBriefings in bioinformatics2026

Influ-BERT: a domain-adaptive genomic language model for advancing influenza A virus research.

Rongye Ye, Lun Li, Ana Tereza Ribeiro de Vasconcelos, Shuhui Song

Abstract read
In one paragraph

Article in Briefings in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Rongye YeNational Genomics Data Center, China National Center for Bioinformation, No. 1 Beichen West Road, Chaoyang District, Beijing 100101, China.ORCID 0009-0001-4456-6319
Lun LiNational Genomics Data Center, China National Center for Bioinformation, No. 1 Beichen West Road, Chaoyang District, Beijing 100101, China.ORCID 0000-0003-3242-031X
Ana Tereza Ribeiro de VasconcelosBioinformatics Laboratory, National Laboratory for Scientific Computing, 333 Getúlio Vargas Avenue, Quitandinha, Petrópolis, RJ 25651‑076, Brazil.
Shuhui SongNational Genomics Data Center, China National Center for Bioinformation, No. 1 Beichen West Road, Chaoyang District, Beijing 100101, China.ORCID 0000-0003-2409-8770

Funding

FAPERJ E-26/201.046/2022 and CNPq (CNPq) 307145/2021-2Key Collaborative Research Program of the Alliance of National and International Science Organizations for the Belt and Road Regions ANSO-CR-KP-2022-09National Key Research & Development Program of China 2023YFC2604400National Key Research & Development Program of China 2024YFC2311303National Key Research & Development Program of China 2025YFF1207901National Natural Science Foundation of China 32270718
6 · The paper itself

Abstract

Influenza A virus (IAV) poses a persistent threat to global public health due to its broad host adaptability, frequent anti-genic variation, and potential for cross-species transmission. Accurate identification of IAV subtypes is essential for effective epidemic surveillance and precise disease control. Here, we present Influ-BERT, a domain-adaptive pretrained model based on the Transformer architecture. Optimized from DNABERT-2, Influ-BERT was developed using a dedicated corpus of ~900 000 influenza genome sequences. We constructed a custom Byte Pair Encoding tokenizer, and employed a two-stage training strategy involving domain-adaptive pretraining followed by task-specific fine-tuning. This approach significantly enhanced identification performance for IAV subtypes. Experimental results demonstrate that Influ-BERT outperforms both traditional machine learning approaches and general genomic language models, such as DNABERT-2, Necleotide Transformer, and MegaDNA, in the task of IAV subtype identification. The model consistently achieved F1-scores above 97% across five subtype classification tasks and exhibited stable performance gains for subtypes that are underrepresented in sequencing data, including H5N8, H1N2, and H13N6. Beyond subtype identification, Influ-BERT was successfully applied to additional tasks including respiratory virus identification, IAV pathogenicity prediction, and identification of IAV genomic fragments and functional genes, demonstrating robust performance throughout. Further interpretability analysis using sliding window perturbation confirmed that the model focuses on biologically significant genomic regions, providing insight into its improved predictive capability.

Indexed as

Genome, ViralGenomicsInfluenza A virusInfluenza, HumanHumansgenomic language modelIAV Pthogenicity PedictionIAV subtype identificationInflu-BERTinfluenza virus

Identifiers

PMID41985060
PMCPMC13082375

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.