Evidence map›Paper›PMID 41993399›Full record

ArticlebioRxiv : the preprint server for biology2026

DrugPlayGround: Benchmarking Large Language Models and Embeddings for Drug Discovery.

Tianyu Liu, Sihan Jiang, Fan Zhang, Kunyang Sun, Teresa Head-Gordon, Hongyu Zhao

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Tianyu LiuInterdepartmental Program in Computational Biology & Bioinformatics, Yale University, USA.
Sihan JiangDepartment of Biostatistics, Yale University, USA.
Fan ZhangDepartment of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong.
Kunyang SunPitzer Theory Center and Department of Chemistry, University of California, Berkeley, Berkeley, CA 94720 USA.
Teresa Head-GordonPitzer Theory Center and Department of Chemistry, University of California, Berkeley, Berkeley, CA 94720 USA.
Hongyu ZhaoInterdepartmental Program in Computational Biology & Bioinformatics, Yale University, USA.ORCID 0000-0003-1195-9607

Funding

Project 5: Pandemic Virus Helicase InhibitorsU19AI171954 · NIAID · UNIVERSITY OF MINNESOTA · PI Reuben S Harris, Fang Li · 2022 to 2026
$100.9M
NIAID NIH HHS U19 AI171954
6 · The paper itself

Abstract

Large language models (LLMs) are in the ascendancy for research in drug discovery, offering unprecedented opportunities to reshape drug research by accelerating hypothesis generation, optimizing candidate prioritization, and enabling more scalable and cost-effective drug discovery pipelines. However there is currently a lack of objective assessments of LLM performance to ascertain their advantages and limitations over traditional drug discovery platforms. To tackle this emergent problem, we have developed DrugPlayGround, a framework to evaluate and benchmark LLM performance for generating meaningful text-based descriptions of physiochemical drug characteristics, drug synergism, drug-protein interactions, and the physiological response to perturbations introduced by drug molecules. Moreover, DrugPlayGround is designed to work with domain experts to provide detailed explanations for justifying the predictions of LLMs, thereby testing LLMs for chemical and biological reasoning capabilities to push their greater use at the frontier of drug discovery at all of its stages.

Indexed as

Drug DiscoveryDrug Target PredictionEmbedding ModelLarge Language ModelPerturbation PredictionSynergy Effect Prediction

Identifiers

PMID41993399
PMCPMC13081941

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.