Evidence map›Paper›PMID 42286684›Full record

ArticleJournal of cheminformatics2026

Benchmarking molecular representations and machine learning algorithms for asymmetric catalysis: a palladium-catalysed decarboxylative asymmetric allylic alkylation case study.

Eduardo Aguilar-Bejarano, Declan Galvin, David M Rogers, Ender Özcan, Simon Woodward, Patrick J Guiry, Grazziela Figueredo

Abstract read
In one paragraph

Article in Journal of cheminformatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Eduardo Aguilar-Bejarano *School of Chemistry, University of Nottingham, University Park, Nottingham, NG7 2RD, UK.
Declan Galvin *School of Chemistry, Centre for Synthesis and Chemical Biology, University College Dublin, BelfieldDublin, D04 N2E5, Ireland.
David M RogersSchool of Chemistry, University of Nottingham, University Park, Nottingham, NG7 2RD, UK.
Ender ÖzcanSchool of Computer Science, University of Nottingham, Jubilee Campus, Nottingham, NG8 1BB, UK.
Simon WoodwardSchool of Chemistry, University of Nottingham, University Park, Nottingham, NG7 2RD, UK. simon.woodward@nottingham.ac.uk.
Patrick J GuirySchool of Chemistry, Centre for Synthesis and Chemical Biology, University College Dublin, BelfieldDublin, D04 N2E5, Ireland. patrick.guiry@ucd.ie.
Grazziela FigueredoSchool of Medicine, University of Nottingham, University Park, NottinghamNottinghamshire, NG7 2RD, UK. g.figueredo@nottingham.ac.uk.

Funding

Engineering and Physical Sciences Research Council 18/EPSRC-CDT/3582Engineering and Physical Sciences Research Council EP/S022236/1Science Foundation Ireland 18/EPSRC-CDT/3582
6 · The paper itself

Abstract

Machine learning (ML) applied to metal-ligand asymmetric catalysis remains less explored compared to other applications, such as drug design or materials science. Current strategies frequently focus on augmenting pre-existing descriptors (e.g. those originally formulated for medicinal chemistry), the development of new bespoke steric and electronic descriptors, and the use of molecular fingerprints. Such method diversity, in the absence of user guidelines, makes selecting optimal ML tools to model new asymmetric catalysis problems challenging. This is exacerbated in early asymmetric catalysis metal-ligand development, where typically only limited data describing how ligands and substrates affect reaction stereoselectivity are available. Herein, we present a benchmarking pipeline for asymmetric catalysis ML studies and a comparative evaluation of reaction representations, including bespoke electronic and steric descriptors, fingerprints (CircuS and Morgan), physicochemical descriptors (from RDKit), machine-learnt representations (ChemBERTa, MolT5, and UniMol), and graph representations, paired with a range of ML algorithms (Support Vector Regressors, Random Forests, Extreme Gradient Boosting, Multilayer Perceptron, and GNNs). We use a database comprising only 103 early-development palladium-catalysed decarboxylative asymmetric allylic alkylation (DAAA) exemplars (three Trost-type ligands and 54 substrates) to predict reaction enantioselectivity, evaluating performance in low-data regimes and in extrapolative tasks under both random and substrate-scaffold splitting. An external validation dataset comprising 19 more recently developed reactions is used as a final evaluation. Across data regimes, Morgan fingerprints, ChemBERTa, and the bespoke V4 descriptors emerged as the most robust representations, with Random Forest the most consistent algorithm; on truly unseen substrates, an ensemble of GNNs gave the most accurate predictions. Importantly, scaffold-based splitting was found to estimate real-world extrapolation performance more reliably than random splitting. Post-hoc explainability analyses (SHAP and integrated gradients) revealed that the bespoke and fingerprint representations provided chemically meaningful insights into the substrate and catalyst features driving enantioselectivity, whereas embedding- and RDKit-based representations did not. Lastly, the methodology was validated on asymmetric hydrogenation and palladium-electrocatalysed C-H activation datasets, demonstrating its applicability to a wide range of asymmetric reactions.Scientific contributionA curated dataset of 103 palladium-catalysed decarboxylative asymmetric allylic alkylation (DAAA) reactions, drawn consistently from Guiry group publications, is introduced together with an external validation set of 19 newer reactions. A bespoke, interpretable descriptor strategy that combines fragment-level steric (van der Waals volume) and electronic (Hammett-derived) parameters is developed for this reaction class, requiring neither DFT computations nor experimental crystal structures. A systematic benchmarking methodology is proposed for evaluating combinations of molecular representations and machine learning algorithms under realistic conditions-including small training sets and chemical-space extrapolation-and is validated across three asymmetric catalysis datasets (DAAA, asymmetric hydrogenation, and C-H activation).

Indexed as

Asymmetric catalysisBenchmarkingExplainable AIMachine learningSubstrate screening

Identifiers

PMID42286684
PMCPMC13495154

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.