Evidence map›Paper›PMID 42380223›Full record

Articlenpj drug discovery2025

Evaluation of DNA encoded library and machine learning model combinations for hit discovery.

Sumaiya Iqbal, Wei Jiang, Eric Hansen, Tonia Aristotelous, Shuang Liu, Andrew Reidenbach, Cerise Raffier, Alison Leed, Chengkuan Chen, Lawrence Chung and 4 more

Abstract read
In one paragraph

Article in npj drug discovery, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 9 papers.

0numbers the graph read from it
0cells of the map it votes in
9citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

9 citing papers in PubMed.

  1. Review
  2. Article
  3. Review
  4. Article
  5. Review
  6. Article
  7. Article
  8. Article
  9. Undersampling techniques for large datasets.bioRxiv : the preprint server for biology · 2025
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Sumaiya IqbalBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA. sumaiya@broadinstitute.org.
Wei JiangBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Eric HansenBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Tonia AristotelousBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Shuang LiuBroad Institute of MIT and Harvard, Chemical Biology and Therapeutics, Cambridge, MA, 02142, USA.
Andrew ReidenbachBroad Institute of MIT and Harvard, Chemical Biology and Therapeutics, Cambridge, MA, 02142, USA.
Cerise RaffierBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Alison LeedBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Chengkuan ChenBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Lawrence ChungBroad Institute of MIT and Harvard, Chemical Biology and Therapeutics, Cambridge, MA, 02142, USA.
Eric SigelSignel IX LLC, Belmont, MA, 02478, USA.
Alex BurginBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Sandy GouldBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA.
Holly H SoutterBroad Institute of MIT and Harvard, Center for the Development of Therapeutics, Cambridge, MA, 02142, USA. hsoutter@broadinstitute.org.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

DNA-Encoded Library (DEL) technology allows the screening of millions to billions of compounds in a pooled fashion, which is faster and cheaper than traditional approaches. The massive amounts of DEL binder and not-binder data enable Machine Learning (ML) model development and virtual screening of readily accessible, drug-like libraries in an ultra-high-throughput fashion. Here, we report a comparative assessment of DEL + ML pipeline for hit discovery using three DELs and five ML models (fifteen DEL + ML combinations). Each ML model was used to identify orthosteric binders of two therapeutic targets, Casein kinase 1α/δ (CK1α/δ). Overall, 10% and 94% of the predicted binders and not-binders were confirmed in biophysical assays, including two nanomolar binders (187 and 69.6 nM). Our study provides insights into the DEL + ML paradigm for hit discovery: the importance of chemical diversity in training data and ML model generalizability over accuracy. We publicly shared our results for further use and similar developments.

Identifiers

PMID42380223
PMCPMC13267125

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.