Evidence map›Paper›PMID 42730098›Full record

ArticleACS omega2026

Machine Learning-Driven Prediction of Dose-Linear Pharmacokinetics: Utilizing Molecular Descriptors to Guide Formulation Strategy.

Elizabeth S Levy, Karen E Samy, Andrea Anelli, Dennis H Leung

Abstract read
In one paragraph

Article in ACS omega, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Elizabeth S LevyDepartment of Synthetic Molecule Pharmaceutical Sciences, Genentech Inc, South San Francisco, California 94080, United States.ORCID https://orcid.org/0000-0001-6802-885X
Karen E SamyDepartment of Drug Metabolism and Pharmacokinetics, Genentech Inc, South San Francisco, California 94080, United States.
Andrea AnelliRoche Pharma Research and Early Development, Therapeutic Modalities, Roche Innovation Center Basel, F. Hoffmann-La Roche Ltd., Basel 4070, Switzerland.
Dennis H LeungDepartment of Synthetic Molecule Pharmaceutical Sciences, Genentech Inc, South San Francisco, California 94080, United States.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Early stage preclinical drug formulation development can be impacted by limited materials and time, especially when making rapid, data-driven decisions is critical. An important parameter to consider is dose linearity, as this can guide the selection between determining if a conventional or enabled formulation strategy is necessary to achieve the expected target exposure as the dose is escalated. This work explores the application of machine learning models to predict dose linearity for compounds administered to the mouse species based on the molecular properties. We utilize a diverse set of models, including Random Forest, XGBoost, LASSO, Neural Networks, and TabPFN, using two distinct data sets for training: a full data set including measured and calculated descriptors and a more accessible data set containing only the RDKit descriptors. Cross-validation was applied with random and scaffold splits, and model performance was evaluated based on a test data set containing 10% of the full data set that was held out during training. Our analysis indicated that the ML models trained with the full descriptors were comparable to the SMILES-only derived descriptors with RDKit. Additionally, while most models achieved comparable performance, LASSO was prioritized due to the interpretability, feature selection, and resistance to overfitting smaller data sets. With a random split and RDKit-only descriptors, the LASSO model resulted in 90.9% precision in the test set and 88.0% accuracy. Precision was optimized to minimize false positives, as incorrectly determining that a compound would not be less-than-dose proportional, where a conventional vehicle would be sufficient, has a higher cost than false negatives. To further evaluate the model's capacity to apply to novel chemical compounds, a scaffold split was tested. A gap in performance was seen with a precision of 89.3% in training and 80.9% in testing, highlighting the challenges in predicting outcomes with new, chemically unique architectures. A significant finding of this work is the robust performance of the models trained solely on accessible RDKit descriptors, demonstrating that a tool can be developed without more resource-intensive measured data. Overall, this research focuses on a data-driven approach based on dose linearity prediction for whether a conventional, nonsolubilizing vehicle is sufficient or if an enabled formulation would be required to achieve the target exposures. The application to forecast a compound's dose linearity risk with only the chemical structures provides a valuable tool to guide early formulation strategy and accelerate compound evaluation and preclinical development during the drug discovery phase.

Identifiers

PMID42730098
PMCPMC13563595

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.