Evidence map›Paper›PMID 42750043›Full record

ArticleGenome biology2026

Genolator enables protein function interpretation using a multimodal large language model fusing genomic and structural interpretation with natural language interaction.

Martin Danner, Tanhim Islam, Matthias Begemann, Florian Kraft, Miriam Elbracht, Ingo Kurth, Jeremias Krause

Abstract read
In one paragraph

Article in Genome biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Martin DannerCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Tanhim IslamCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Matthias BegemannCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Florian KraftCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Miriam ElbrachtCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Ingo KurthCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany.
Jeremias KrauseCenter for Human Genetics and Genomic Medicine, Medical Faculty, Uniklinik RWTH Aachen, Pauwelsstrasse 30, D-52074, Aachen, Germany. jerkrause@ukaachen.de.ORCID https://orcid.org/0000-0001-9915-7400

Funding

Medizinische Fakultät, RWTH Aachen University START
6 · The paper itself

Abstract

backgroundDecoding the genetic code to unveil its genome functionality is a monumental task which would greatly advance the understanding of disease mechanisms and development of targeted treatments. Although large language models (LLMs) have transformed natural language processing across diverse domains, translating the complex language of DNA into human-readable form remains challenging due to genomic data complexity and unexplored regions of the human genome. Current language models either are capable of processing natural language or the genomic code. Models fusing both aspects are largely lacking.

resultsHere we present Genolator, a multimodal large language model that integrates embeddings from DNA sequences, amino acid sequences, and protein structures with natural language queries. Fine-tuned on over 365,000 question-answer pairs generated using abstracted Gene-Ontology (GO) terms, Genolator effectively answers queries regarding protein subcellular localization, molecular function, and biological processes. Evaluation demonstrates high accuracy in confirming or denying protein function associations, outperforming baseline models such as openly available allrounder LLMs like GPT 4.1 as well as smaller domain-specific models integrating knowledge from a protein structure transformer. Explorations of Genolator's hidden states unveil a biologically and linguistically plausible organization of its learned representations. Analysis of the attention heads of the underlying language model and an ablation study provide evidence for a benefit of the multi-modal approach.

conclusionGenolator enhances accessibility to genomic information by enabling natural language interaction with protein data, facilitating biological discovery, and clinical research. It represents a step towards bridging genomic code and human language through the integration of a multimodal LLM.

Indexed as

GenomicsNatural Language ProcessingProteinsSoftwareHumansLarge Language ModelsProteinsFeature fusionGenolatorGenomic language modelsMultimodal language modelsProtein function prediction

Identifiers

PMID42750043
PMCPMC13579901

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.