Evidence map›Paper›PMID 37732243›Full record

ArticlebioRxiv : the preprint server for biology2023

Poor Generalization by Current Deep Learning Models for Predicting Binding Affinities of Kinase Inhibitors.

Wern Juin Gabriel Ong, Palani Kirubakaran, John Karanicolas

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Wern Juin Gabriel OngCancer Signaling & Microenvironment Program, Fox Chase Cancer Center, Philadelphia, PA 19111.
Palani KirubakaranCancer Signaling & Microenvironment Program, Fox Chase Cancer Center, Philadelphia, PA 19111.
John KaranicolasCancer Signaling & Microenvironment Program, Fox Chase Cancer Center, Philadelphia, PA 19111.

Funding

WORD PROCESSING CENTER--COREP30CA006927 · NCI · RESEARCH INST OF FOX CHASE CAN CTR · PI Eric Andrew Ross · 1985 to 2026
$138.8M
Designing selective kinase inhibitors via deep learningR01GM141513 · NIGMS · RESEARCH INST OF FOX CHASE CAN CTR · PI RINK, LORI · 2022 to 2025
$2.4M
NCI NIH HHS P30 CA006927NIGMS NIH HHS R01 GM141513
6 · The paper itself

Abstract

The extreme surge of interest over the past decade surrounding the use of neural networks has inspired many groups to deploy them for predicting binding affinities of drug-like molecules to their receptors. A model that can accurately make such predictions has the potential to screen large chemical libraries and help streamline the drug discovery process. However, despite reports of models that accurately predict quantitative inhibition using protein kinase sequences and inhibitors' SMILES strings, it is still unclear whether these models can generalize to previously unseen data. Here, we build a Convolutional Neural Network (CNN) analogous to those previously reported and evaluate the model over four datasets commonly used for inhibitor/kinase predictions. We find that the model performs comparably to those previously reported, provided that the individual data points are randomly split between the training set and the test set. However, model performance is dramatically deteriorated when all data for a given inhibitor is placed together in the same training/testing fold, implying that information leakage underlies the models' performance. Through comparison to simple models in which the SMILES strings are tokenized, or in which test set predictions are simply copied from the closest training set data points, we demonstrate that there is essentially no generalization whatsoever in this model. In other words, the model has not learned anything about molecular interactions, and does not provide any benefit over much simpler and more transparent models. These observations strongly point to the need for richer structure-based encodings, to obtain useful prospective predictions of not-yet-synthesized candidate inhibitors.

Identifiers

PMID37732243
PMCPMC10508770

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.