Evidence map›Paper›PMID 41421504›Full record

ReviewAdvanced drug delivery reviews2026

Small data, big challenges: Machine- and deep-learning strategies for data-limited drug discovery.

Nazreen Pallikkavaliyaveetil, Sriram Chandrasekaran

Abstract readReview
In one paragraph

Review in Advanced drug delivery reviews, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.

0numbers the graph read from it
0cells of the map it votes in
14citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

14 citing papers in PubMed.

  1. HRSC advances · 2026
    Article
  2. Review
  3. Review
  4. Article
  5. Review
  6. Article
  7. Review
  8. Review
  9. Review
  10. Review
  11. Review
  12. Review
  13. Review
  14. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

2 authors.

Nazreen PallikkavaliyaveetilDepartment of Biomedical Engineering, University of Michigan, Ann Arbor, MI 48109, USA; Michigan Institute for Data & AI in Society (MIDAS), University of Michigan, Ann Arbor, MI 48109, USA.
Sriram ChandrasekaranDepartment of Biomedical Engineering, University of Michigan, Ann Arbor, MI 48109, USA; Center for Bioinformatics and Computational Medicine, Ann Arbor, MI 48109, USA; Program in Chemical Biology, University of Michigan, Ann Arbor, MI 48109, USA; Rogel Cancer Center, University of Michigan Medical School, Ann Arbor, MI 48109, USA; Michigan Institute for Data & AI in Society (MIDAS), University of Michigan, Ann Arbor, MI 48109, USA. Electronic address: csriram@umich.edu.

Funding

Linking metabolic activity with drug sensitivity using metabolic influence networksR35GM137795 · NIGMS · UNIVERSITY OF MICHIGAN AT ANN ARBOR · PI Sriram Chandrasekaran · 2020 to 2026
$2.6M
NIGMS NIH HHS R35 GM137795
6 · The paper itself

Abstract

A critical bottleneck limiting the potential of Machine Learning (ML) and Deep Learning (DL) models within the drug discovery and development (DDD) pipeline is the scarcity of high-quality experimental data. Limited data is not an anomaly but an inherent characteristic of the DDD process. Significant financial costs, time, and confidentiality concerns limit the scale of available datasets. Applying standard ML and DL algorithms directly to these small datasets presents substantial challenges. Traditional ML models remain constrained by their dependence on handcrafted features and limited ability to capture complex biological relationships. In contrast, DL algorithms that assume data abundance are prone to overfitting and poor generalization when trained on small datasets. The small data problem thus represents a fundamental constraint that shapes the practical utility and trustworthiness of AI applications in DDD. While prior reviews have surveyed the broad landscape of AI and ML in drug discovery, a significant gap exists concerning the small data challenge across the DDD pipeline. Addressing this challenge requires adapting DL methods that typically assume data abundance, while also extending traditional ML approaches that, although well-suited to small data, remain limited in their representational capacity. This review addresses this gap by surveying key drug discovery tasks, highlighting the prevalence of limited data, and synthesizing both traditional ML methods and advanced DL strategies tailored to these contexts. By integrating methodological advances with task-specific applications, the review outlines current approaches and identifies opportunities for advancing robust, interpretable, and generalizable AI in drug discovery.

Indexed as

Deep LearningDrug DiscoveryMachine LearningAlgorithmsBig DataHumansData augmentationDeep learningFew-shot learningLimited-data drug discoveryMachine learningSelf-supervised learningTransfer learning

Identifiers

PMID41421504
PMCPMC13094559

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.