Evidence map›Paper›PMID 42794906›Full record

ReviewCancers2026

Big Data and Artificial Intelligence in Cancer Drug Discovery: Promise, Challenges, and Emerging Opportunities.

Fakhar U Singhera, Justin M Overhulse, Terrence M Lee, Jonathan E Katz, Jerry S H Lee, Charles E McKenna

Abstract readReview
In one paragraph

Review in Cancers, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Fakhar U SingheraMork Family Department of Chemical Engineering and Materials Science, University of Southern California, Los Angeles, CA 90089, USA.ORCID 0009-0000-8681-5769
Justin M OverhulseDepartment of Chemistry, University of Southern California, Los Angeles, CA 90089, USA.ORCID 0000-0002-1902-9079
Terrence M LeeDepartment of Chemistry, University of Southern California, Los Angeles, CA 90089, USA.
Jonathan E KatzEllison Medical Institute, Los Angeles, CA 90064, USA.ORCID 0000-0003-2699-9708
Jerry S H LeeMork Family Department of Chemical Engineering and Materials Science, University of Southern California, Los Angeles, CA 90089, USA.ORCID 0000-0003-1515-0952
Charles E McKennaDepartment of Chemistry, University of Southern California, Los Angeles, CA 90089, USA.ORCID 0000-0002-3540-6663

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Oncologic drug development is lengthy (~14 years) and expensive (~1.2 billion USD) with low clinical trial success rates (4.1%). Big data and artificial intelligence (AI) are widely proposed as tools to address these challenges. In this review, we examine the current performance and future potential of big data and AI applied to preclinical discovery and development, clinical trials, and the regulatory approval process. We first examine the data foundation required for effective AI, including data harmonization, data commons, and analytical tools. We then assess preclinical applications spanning target identification, compound-library curation, virtual ligand screening, generative chemical design, and high-throughput and high-content screening. In clinical development, we consider the use of big data and AI for outcome prediction, trial design, external and synthetic control arms, adaptive monitoring, and in silico trials. Finally, we discuss how post-approval electronic health records can generate real-world data and real-world evidence to support drug repurposing and improve future oncology drug discovery. Big data is conventionally characterized by a series of "Vs." In this review, we have used seven "Vs" spanning descriptive and constraining properties of big data and a singular outcome. We have proposed an eighth, Vernacular, a constraint defined as the combined alignment of data semantics and terminology, data representation, data exchange, and data governance across heterogeneous, independently generated datasets to promote interoperability and combined analysis. Although cancer data exhibit substantial Volume, Velocity, and Variety, they remain distributed across fragmented repositories that often cannot be readily integrated. We conclude with a discussion of tabulated resources currently available for the application of big data and AI to oncologic therapeutics.

Indexed as

artificial intelligencebig datacancerclinical trialsdrug discoveryhigh content screeninghigh throughput screeningmachine learningreal world datareal world evidencevirtual ligand screening

Identifiers

PMID42794906
PMCPMC13604657

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.