Evidence map›Paper›PMID 38948789›Full record

ArticlebioRxiv : the preprint server for biology2024

Signals in the Cells: Multimodal and Contextualized Machine Learning Foundations for Therapeutics.

Alejandro Velez-Arce, Michelle M Li, Wenhao Gao, Xiang Lin, Kexin Huang, Tianfan Fu, Bradley L Pentelute, Manolis Kellis, Marinka Zitnik

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Alejandro Velez-ArceDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA 02115.ORCID 0009-0009-2303-6114
Michelle M LiDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA 02115.ORCID 0000-0003-0223-7485
Wenhao GaoDepartment of Chemical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139.
Xiang LinDepartment of Biomedical Informatics, Harvard Medical School, Boston, MA 02115.
Kexin HuangDepartment of Computer Science, Stanford School of Engineering, Stanford, CA 94305.
Tianfan FuDepartment of Computational Science, Rensselaer Polytechnic Institute, Troy, NY 12180.
Bradley L PenteluteDepartment of Chemistry, Massachusetts Institute of Technology, Cambridge, MA 02139.
Manolis KellisBroad Institute of MIT and Harvard, Computer Science and Artificial Intelligence Laboratory, MIT, Electrical Engineering and Computer Science Department, Massachusetts Institute of Technology, Cambridge, MA 02139.ORCID 0000-0001-7113-9630
Marinka ZitnikBroad Institute of MIT and Harvard, Harvard Data Science Initiative, Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Department of Biomedical Informatics, Harvard Medical School, Boston, MA 02215.ORCID 0000-0001-8530-7228

Funding

Training Program in Bioinformatics and Integrative GenomicsT32HG002295 · NHGRI · MASSACHUSETTS INSTITUTE OF TECHNOLOGY · PI Peter J Park · 2001 to 2026
$15.8M
Measuring Neonatal RegionalizationR01HD108794 · NICHD · STANFORD UNIVERSITY · PI Jochen Profit, JEANNETTE A ROGOWSKI · 2023 to 2026
$2.8M
NHGRI NIH HHS T32 HG002295NICHD NIH HHS R01 HD108794
6 · The paper itself

Abstract

Drug discovery AI datasets and benchmarks have not traditionally included single-cell analysis biomarkers. While benchmarking efforts in single-cell analysis have recently released collections of single-cell tasks, they have yet to comprehensively release datasets, models, and benchmarks that integrate a broad range of therapeutic discovery tasks with cell-type-specific biomarkers. Therapeutics Commons (TDC-2) presents datasets, tools, models, and benchmarks integrating cell-type-specific contextual features with ML tasks across therapeutics. We present four tasks for contextual learning at single-cell resolution: drug-target nomination, genetic perturbation response prediction, chemical perturbation response prediction, and protein-peptide interaction prediction. We introduce datasets, models, and benchmarks for these four tasks. Finally, we detail the advancements and challenges in machine learning and biology that drove the implementation of TDC-2 and how they are reflected in its architecture, datasets and benchmarks, and foundation model tooling.

Identifiers

PMID38948789
PMCPMC11212894

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.