Evidence map›Paper›PMID 41511463›Full record

SynthesisBriefings in bioinformatics2026

A systematic review of molecular representation learning foundation models.

Bosheng Song, Jiayi Zhang, Ying Liu, Yuansheng Liu, Jing Jiang, Sisi Yuan, Xia Zhen, Yiping Liu

Abstract readSystematic Review
In one paragraph

Synthesis in Briefings in bioinformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Review
  3. Review
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Bosheng SongCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.
Jiayi ZhangCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.
Ying LiuCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.
Yuansheng LiuCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.
Jing JiangCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.
Sisi YuanSchool of Chinese Medicine, Hong Kong Baptist University, 15 Baptist University Road, Kowloon Tong, Kowloon, Hong Kong SAR 999077, China.
Xia ZhenNational Laboratory for Parallel and Distributed Processing, School of Computer, National University of Defense Technology, No. 109 Deya Road, Kaifu District, 410086 Changsha, China.
Yiping LiuCollege of Computer Science and Electronic Engineering, Hunan University, 116 Lushan South Road, Yuelu District, 410086 Changsha, China.ORCID 0000-0001-7340-2551

Funding

Hunan Provincial Natural Science Foundation of China 2024JJ4015National Natural Science Foundation of China 62202153National Natural Science Foundation of China 62272151National Natural Science Foundation of China 62472152National Natural Science Foundation of China 62522110
6 · The paper itself

Abstract

Molecular representation learning (MRL) is afoundation in leveraging computational methods for drug discovery, enabling the transformation of molecular structure and properties into numerical vectors. These vectors serve as input for machine learning models and facilitate the prediction and analysis of molecular attributes, functions, and reactions. The advent of foundation models has introduced both new opportunities and challenges to MRL. These models have improved generalizability and migration in scarce data. Through pretraining and fine-tuning, foundation models can be adapted to various domains. Their robust encoding and generative abilities also allow the transformation of molecular data into more expressive forms. This paper provides a detailed review of current mainstream molecular descriptors and datasets, focusing primarily on the representation of small molecules while excluding larger molecules such as proteins and peptides. It classifies foundation models into two primary categories based on the form of input: unimodal-based and multimodal-based models. For each category, representative models are identified and their advantages and disadvantages evaluated. Moreover, we systematically summarize four core pretraining strategies for MRL foundation models, analyzing their task designs, applicable scenarios, and impacts on downstream performance. In addition, the application of molecular representation foundation models in drug discovery and development is discussed, together with the current status of model interpretability. The paper concludes with insights into the future directions of MRL foundation models.

Indexed as

Computational BiologyDrug DiscoveryMachine LearningModels, MolecularHumansRepresentation Machine Learningdrug discoveryfoundation modelsmachine learningmolecular representation learning

Identifiers

PMID41511463
PMCPMC12784970

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.