Evidence map›Paper›PMID 42415674›Full record

ArticleJournal of chemical information and modeling2026

Modeling the Sensitivity of Large-Scale Virtual Screening to Scoring Function Accuracy, Artifacts, and Library Composition.

Laust Moesgaard, Brian K Shoichet, Olivier Mailhot

Abstract read
In one paragraph

Article in Journal of chemical information and modeling, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Laust MoesgaardDepartment of Physics, Chemistry and Pharmacy, University of Southern Denmark, Odense M5230Denmark.ORCID 0000-0002-5509-9832
Brian K ShoichetDepartment of Pharmaceutical Chemistry, University of California, San Francisco, San Francisco, California94143-2550, United States.ORCID 0000-0002-6098-7367
Olivier MailhotInstitute for Research in Immunology and Cancer, Faculty of Pharmacy, Université de Montréal, Montréal, QuebecH3T 1N8, Canada.

Funding

Development and Testing of New Computational Methods for Ligand Discovery and MechanismR35GM122481 · NIGMS · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI Brian K Shoichet · 2017 to 2026
$8.5M
Rational prioritization algorithm for docking-based virtual screening of trillion-scale make-on-demand small molecule librariesF32GM154469 · NIGMS · UNIVERSITY OF CALIFORNIA, SAN FRANCISCO · PI MAILHOT, OLIVIER · 2024 to 2025
$149k
NIGMS NIH HHS F32 GM154469NIGMS NIH HHS R35 GM122481
6 · The paper itself

Abstract

Large library docking has emerged as a productive approach for ligand discovery, yet a quantitative framework for understanding how docking performance responds to methodological improvements has been lacking. Here, we develop such a framework by modeling large-scale experiments from three previously published docking campaigns, in which 2,682 ligands had been synthesized and tested across the scoring landscape (poor scores, mediocre scores, high scores). The observed experimental hit-rate curves can be reproduced by a simple bivariate normal distribution model, where docking score is interpreted as a noisy predictor of binding free energy. To account for the plateauing and subsequent drop in hit rates often seen at highly favorable docking scores, we add a term for high-ranking docking artifacts, a phenomenon we observe across targets. From this model, three predictions about the sensitivity of docking performance emerge. First, even slight improvements in scoring accuracy would substantially improve both hit rates and hit affinities: quantitatively, a 0.1 increase in the correlation between docking score and binding affinity would justify accepting a ∼10-fold increase in computational cost per molecule, arguing for reinvestment in scoring function accuracy in library docking. Second, docking artifacts, while hard to anticipate, can come to dominate top-scoring lists as libraries grow. Physically testing molecules across a range of log-normalized ranks (pProp) is therefore essential to identify the peak hit rate for a given campaign. Third, prefiltering a library to enrich for molecules with appropriate physicochemical features increases the intrinsic hit rate and substantially boosts docking performance, particularly at tera-scale, with effects comparable to a meaningful improvement in scoring accuracy. Beyond docking, the model's parameters (affinity distribution, score-affinity correlation, artifact frequency) can be fit to any screening method with sufficient experimental data, providing an objective basis for benchmarking and comparing virtual screening approaches. These findings offer a practical framework for optimizing large-scale virtual screening as chemical libraries continue to grow.

Indexed as

ArtifactsMolecular Docking SimulationSmall Molecule LibrariesDrug Evaluation, PreclinicalLigandsProtein BindingProteinsThermodynamicsLigandsProteinsSmall Molecule Libraries

Identifiers

PMID42415674
PMCPMC13418168

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.