Evidence map›Paper›PMID 42123473›Full record

ReviewInternational journal of molecular sciences2026

Integrating Protein Language Models with Multimodal Embeddings to Accelerate Function Prediction of Uncharacterized Proteins.

Ruyang Cheng, Tianyu Liu, Chentao Liao, Xiaomin Wu, Lingyun Zhu, Shaowei Zhang

Abstract readReview
In one paragraph

Review in International journal of molecular sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Review
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Ruyang ChengCollege of Science, National University of Defense Technology, Changsha 410073, China.
Tianyu LiuCollege of Science, National University of Defense Technology, Changsha 410073, China.
Chentao LiaoCollege of Science, National University of Defense Technology, Changsha 410073, China.
Xiaomin WuCollege of Science, National University of Defense Technology, Changsha 410073, China.
Lingyun ZhuCollege of Science, National University of Defense Technology, Changsha 410073, China.
Shaowei ZhangCollege of Science, National University of Defense Technology, Changsha 410073, China.ORCID 0000-0001-9763-3266

Funding

Hunan Province Science and Technology Innovation Program 2024RC3144National Natural Science Foundation of China No. 32401056
6 · The paper itself

Abstract

Accurate prediction of protein function is fundamental to progress in biotechnology and biomedicine, yet progress remains severely hampered by the widening chasm between exponentially growing genomic data and the limited capacity for functional annotation. High-throughput sequencing and metagenomics have driven an explosion in sequence data that far outstrips experimental characterization. UniProt now contains over 203 million protein entries, of which only ~2% have been experimentally validated. This widening "sequence-function gap" exceeds the reach of traditional homology-based tools such as BLAST (v2.17.0) and HMMER (v3.2), which are inherently constrained by sequence identity thresholds. The emergence of Protein Language Models (PLMs), including ESM and ProtTrans, has introduced a transformative paradigm, thereby shifting functional inference from similarity-based retrieval to geometric reasoning within learned semantic spaces. Nevertheless, current approaches remain largely confined to unimodal or narrowly bimodal frameworks, failing to capture the inherently multidimensional determinants of enzymatic function, including active-site geometry, chemical reaction logic, and literature-embedded semantic context. This review systematically adopts a multimodal global-fusion perspective, elucidating how three-dimensional geometric features, chemical reaction semantics, and textual knowledge graphs are synergistically integrated around PLMs as a core backbone. We delineate complementary mechanisms and integration strategies that together enable fine-grained protein function annotation beyond the performance ceiling of single-sequence methods. Furthermore, we survey the translational potential of such frameworks from computational prediction to real biological applications, and critically examine persistent bottlenecks including activity cliffs, transition-state inference, and conformational dynamics. We identify the integration of physics-informed machine learning with dynamics-aware architectures as a pivotal direction toward a causal, mechanism-level understanding of protein function.

Indexed as

Computational BiologyProteinsDatabases, ProteinLarge Language ModelsProteinschemical reaction spacemultimodalprotein function predictionprotein language modelsprotein structure–function

Identifiers

PMID42123473
PMCPMC13164483

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.