Evidence map›Paper›PMID 41258076›Full record

ReviewExperimental & molecular medicine2026

A survey on large language models in biology and chemistry.

Islambek Ashyrmamatov, Su Ji Gwak, Su-Young Jin, Ikhyeong Jun, Umit V Ucak, Jay-Yoon Lee, Juyong Lee

Abstract readReview
In one paragraph

Review in Experimental & molecular medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Islambek Ashyrmamatov *Research Institute of Pharmaceutical Science, College of Pharmacy, Seoul National University, Seoul, Republic of Korea.
Su Ji Gwak *Graduate School of Data Science, Seoul National University, Seoul, Republic of Korea.
Su-Young JinGraduate School of Data Science, Seoul National University, Seoul, Republic of Korea.
Ikhyeong JunDepartment of Molecular Medicine and Biopharmaceutical Sciences, Graduate School of Convergence Science and Technology and College of Pharmacy, Seoul National University, Seoul, Republic of Korea.
Umit V UcakResearch Institute of Pharmaceutical Science, College of Pharmacy, Seoul National University, Seoul, Republic of Korea. braket@snu.ac.kr.ORCID http://orcid.org/0000-0002-9088-0915
Jay-Yoon LeeGraduate School of Data Science, Seoul National University, Seoul, Republic of Korea. lee.jayyoon@snu.ac.kr.
Juyong LeeResearch Institute of Pharmaceutical Science, College of Pharmacy, Seoul National University, Seoul, Republic of Korea. nicole23@snu.ac.kr.ORCID http://orcid.org/0000-0003-1174-4358

Funding

MOE | Korea Environmental Industry and Technology Institute (KEITI) RS-2023-00219144National Research Foundation of Korea (NRF) 2020M3A9G7103933National Research Foundation of Korea (NRF) 2022M3E5F3081268National Research Foundation of Korea (NRF) 2022R1C1C1005080National Research Foundation of Korea (NRF) RS-2023-00256320
6 · The paper itself

Abstract

Artificial intelligence (AI) is reshaping biomedical research by providing scalable computational frameworks suited to the complexity of biological systems. Central to this revolution are bio/chemical language models, including large language models, which are reconceptualizing molecular structures as a form of 'language' amenable to advanced computational techniques. Here we critically examine the role of these models in biology and chemistry, tracing their evolution from molecular representation to molecular generation and optimization. This review covers key molecular representation strategies for both biological macromolecules and small organic compounds-ranging from protein and nucleotide sequences to single-cell data, string-based chemical formats, graph-based encodings and three-dimensional point clouds-highlighting their respective advantages and inherent limitations in AI applications. The discussion further explores core model architectures, such as bidirectional encoder representations from transformers-like encoders, generative pretrained transformer-like decoders and encoder-decoder transformers, alongside their sophisticated pretraining strategies such as self-supervised learning, multitask learning and retrieval-augmented generation. Key biomedical applications, spanning protein structure and function prediction, de novo protein design, genomic analysis, molecular property prediction, de novo molecular design, reaction prediction and retrosynthesis, are explored through representative studies and emerging trends. Finally, the review considers the emerging landscape of agentic and interactive AI systems, showcasing briefly their potential to automate and accelerate scientific discovery while addressing critical technical, ethical and regulatory considerations that will shape the future trajectory of AI in biomedicine.

Indexed as

Artificial IntelligenceComputational BiologyHumansLarge Language Models

Identifiers

PMID41258076
PMCPMC13144691

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.