Evidence map›Paper›PMID 42431910›Full record

ArticleNature communications2026

Navigating chemical-linguistic sharing space with heterogeneous molecular encoding.

Liuzhenghao Lv, Hao Li, Yu Wang, Zijun Chen, Zhiyuan Yan, Zongying Lin, Yuyang Liu, Li Yuan, Yonghong Tian

Abstract read
In one paragraph

Article in Nature communications, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Generative Deep Learning for de Novo Drug Design─A Chemical Space Odyssey.Journal of chemical information and modeling · 2025
    Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Liuzhenghao Lv *Beijing Key Laboratory of Brain-inspired Spiking Large Models, School of Computer Science, Peking University, Beijing, China.
Hao Li *Guangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China.
Yu WangBeijing Key Laboratory of Brain-inspired Spiking Large Models, School of Computer Science, Peking University, Beijing, China.
Zijun ChenGuangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China.
Zhiyuan YanGuangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China.
Zongying LinGuangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China.ORCID http://orcid.org/0009-0001-2035-1970
Yuyang LiuGuangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China.
Li YuanGuangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University, Shenzhen, China. yuanli-ece@pku.edu.cn.ORCID http://orcid.org/0000-0002-2120-5588
Yonghong TianBeijing Key Laboratory of Brain-inspired Spiking Large Models, School of Computer Science, Peking University, Beijing, China. yhtian@pku.edu.cn.ORCID http://orcid.org/0000-0002-2978-5935

Funding

China Postdoctoral Science Foundation 2024M760113China Postdoctoral Science Foundation BX20240013National Natural Science Foundation of China (National Science Foundation of China) 62425101
6 · The paper itself

Abstract

Chemical language models are powerful tools for navigating chemical space, but their reliance on linear representations such as molecular strings creates a semantic gap, hindering their ability to bridge natural language with the full complexity of molecular structures. Here we show that chemical language models can gain a comprehensive, multi-modal understanding of molecules through heterogeneous molecular encoding, which integrates one-dimensional sequences, two-dimensional topology, three-dimensional geometry, and statistically derived molecular fragments. We further introduce a query-based module that converts heterogeneous structural information into a unified representation compatible with language models, together with a chain-of-fragment mechanism that guides molecular generation through a hierarchical chemical blueprinting process. To support research in this area, we constructed a million-scale dataset for multi-objective molecular design. Experimentally, the framework enables bidirectional navigation of the chemical-linguistic space, achieving consistent improvements across molecular comprehension and design tasks over strong baselines.

Identifiers

PMID42431910
PMCPMC13482991

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.