Data Sets, Tools & Workflows

Data sets, tools and workflows

C-CAS is committed to making all data sets, tools and workflows freely available to the scientific community. This page provides a one-stop list of C-CAS contributions to the field of data chemistry, together with the publications and, where available, tutorial material developed within C-CAS. In addition, the DOI links to datasets are given on the publication pages. If you use the data sets and tools, please ensure to cite the appropriate publications and DOIs.


Open Reaction Database

The Open Reaction Database (ORD) is an open-access schema and infrastructure for structuring and sharing organic reaction data, including a centralized data repository. The ORD schema supports conventional and emerging technologies, from benchtop reactions to automated high-throughput experiments and flow chemistry.

Open Reaction Database

Publications:

Kearnes SM, Maser MR, Wleklinski M, Kast A, Doyle AG, Dreher SD, Hawkins JM, Jensen KF, Coley CW. The Open Reaction Database. J. Am. Chem. Soc. 2021, 143, 18820-18826. doi.org/10.1021/jacs.1c09820

Mercado R, Kearnes SM, Coley C. Data Sharing in Chemistry: Lessons Learned and a Case for Mandating Structured Reaction Data. J. Chem. Inf. Model. 2023, 63, 4253-4265. doi.org/10.1021/acs.jcim.3c00607

Tutorial material:

ORD Tutorial


Multi-Linear Regression

Multilinear regressions (MLR) is a widely used tool in predictive chemistry and the elucidation of mechanism. C-CAS researchers developed a variety of tools for conducting and visualizing MLR analyses, the rapid generation of chemically relevant features, and feature databases such as Kraken (see above).

The different components of the MLR workflows are:

1) Workflow scripts for automated collection of molecular properties, as well as atom- and bond-level properties for a conserved moiety of interest from Gaussian jobs. Post-processing allows for the collection of condensed descriptors for conformational ensembles. https://doi.org/10.5281/zenodo.19560208

2) Open-source Python package for generation of SMART probe ensembles and calculation of SMART molecular descriptors https://doi.org/10.5281/zenodo.19560206

3) Workflow for linear modeling, primarily driven by bidirectional stepwise MLR, as well as threshold analysis and tools for feature curation https://doi.org/10.5281/zenodo.19560132

4) Workflow for classification of chemical compounds as active or inactive based on experimental outputs and a set of previously computed descriptors based on sci-kit learn's DecisionTreeClassifier. It yields models that can be both interpretable and predictive. https://doi.org/10.5281/zenodo.19560178

The analysis section also includes scripts for using HoloViews (Bokeh backend) to generate interactive plots https://doi.org/10.5281/zenodo.19560202


Kraken

Kraken is a discovery platform covering monodentate organophosphorus(III) ligands providing comprehensive physicochemical descriptors based on representative conformer ensembles. Using quantum-mechanical methods, we calculated descriptors for 1558 ligands, including commercially available examples, and trained machine learning models to predict properties of over 300000 new ligands.

Kraken

The molecular descriptors in Kraken form the basis of the Phosphine Predictor, a web tool by Sigma-Aldrich for the selection of phosphine ligands for cross-coupling reactions.

Publications:

Gensch, T.; dos Passos Gomes, G.; Friederich, P.; Peters, E.; Gaudin, T.; Pollice, R.; Jorner, K.; Nigam, A.; Lindner-D'Addario, M.; Sigman, M. S.; Aspuru-Guzik, A. A Comprehensive Discovery Platform for Organophosphorus Ligands for Catalysis. J. Am. Chem. Soc. 2022, 144, 3, 1205–1217. doi.org/10.1021/jacs.1c09718

Tutorial Material:

Kraken Tutorial


Auto-QChem

Auto-QChem is an automatic, high-throughput and end-to-end DFT calculation workflow that computes chemical descriptors for organic molecules. Tailored toward users without extensive programming experience, Auto-QChem has facilitated more than 38 000 DFT calculations for 17 000 molecules as of January 2022. Starting from string representations of molecules, Auto-QChem automatically (a) generates conformational ensembles, (b) submits and manages DFT calculations on a high-performance computing (HPC) cluster, (c) extracts production-ready features that are suitable for statistical analysis and machine learning model development, and (d) stores resulting calculations in a cloud-hosted and web-accessible database

Auto-QChem Github 

Auto-QChem Web Interface

Publication:

Żurański, A.M.; Wang, J.Y.; Shields, B.J.; Doyle, A.G. Auto-QChem: an automated workflow for the generation and storage of DFT calculations for organic molecules. React. Chem. Engin. 2022, 7, 1276-1284 doi.org/10.1039/D2RE00030J


DBStep

DBStep is a python package for obtaining DFT-Based Steric Parameters from 3-dimensional chemical structures. It can parse the outputs from most computational chemistry programs and other common molecular structure file formats. Steric properties can either be obtained exactly or by using a Cartesian grid, the latter approach being amenable to the featurization of a molecular isodensity surface (DBSTEP can process wavefunction files) rather than using classical atomic radii. Currently, traditional Sterimol parameters (L, Bmin, Bmax) and percent buried volume parameters are implemented, as well as our novel steric parameter vectors Sterimol2vec and vol2vec. This package is designed for use on the command line or alternatively implemented in a Python script for use in a computational workflow to collect steric parameters

DB Step Download

Publication:

Luchini, G.; Patterson, T.; Paton, R. S. DBSTEP: DFT Based Steric Parameters. 2022, DOI: 10.5281/zenodo.4702097 

Tutorial Material:

DB Step Tutorial

Overview of modern steric parameters


Cascade

CASCADE stands for ChemicAl Shift CAlculation with DEep learning. It is a stereochemistry-aware online calculator for NMR chemical shifts using a graph network approach developed at Colorado State University. Molecular input can be specified as SMILES or through the graphical interface. An automated workflow executes 3D structure embedding and MMFF conformer searching. The full ensemble of optimized conformations are passed to a trained graph neural network to predict the NMR chemical shift (in ppm) for each carbon.

Cascade Download

Cascade Web Interface

Publication:

Guan, Y.; Sowndarya, S. V. S.; Gallegos, L. C.; St. John P. C.; Paton, R. S. Real-time prediction of 1H and 13C chemical shifts with DFT accuracy using a 3D graph neural network Chem. Sci. 2021, 12, 12012-12026 doi.org/10.1039/D1SC03343C


AQME

AQME is an ensemble of automated QM workflows, including: 1) RDKit- and CREST-based conformer generator and ready-to-submit QM input files starting from individual files or databases, 2) post-processing of QM output files to fix extra imaginary frequencies, unfinished jobs and error terminate

AQME

Publication:

Alegre-Requena, J. V.; Sowndarya, S.; Pérez-Soto, R.; Alturaifi, T.; Paton, R. AQME: Automated Quantum Mechanical Environments for Researchers and Educators. Wiley Interdiscip. Rev. Comput. Mol. Sci. 2023, 13, e1663 doi.org/10.1002/wcms.1663

Tutorial Material:

AQME Tutorial


EDBO+

EDBO+ is a multi-objective reaction Bayesian optimization platform that builds on the previously published Bayesian optimizers EDBO. The web-based application incorporates features such as condition modification on the fly and data visualization.

EDBO+ Web Interface

 EDBO+ Github

Publications:

EDBO+ Torres, J. A. G.; Lau, S. H.; Anchuri, P.; Stevens, J. M; Tabora, J. E; Li, J.; Borovika, A.; Adams, R. P; Doyle, A.G. A Multi-Objective Active Learning Platform and Web App for Reaction Optimization. J. Am. Chem. Soc. 2022, 144, 1999-2007. doi/10.1021/jacs.2c08592

EDBO Shields, B.J.; Stevens, J.; Li, J.; Prarasram, M.; Damani, F.; Martinez Alvaro, J., Janey, J. Adams, R.P., Doyle, A. Bayesian Reaction Optimization as A Tool for Chemical Synthesis. Nature 2021, 590, 89-96. doi.org/10.1038/s41586-021-03213-y

Tutorial material:

EDBO+ Tutorial

Intro to Bayesian Optimization

"Over the Arrow’ Optimization


Bandit-Optimization

This code uses reinforcement learning, specifically the multi-armed bandit approach, for reaction optimization. It demonstrates data-efficient learning at high accuracies and has unique functionalities.

Bandit Optimizer Code.

Publication:

Wang JY, Stevens JM, Kariofillis SK, Tom MJ, Golden DL, Li J, Tabora JE, Parasram M, Shields BJ, Primer DN, Hao B. Identifying general reaction conditions by bandit optimization. Nature. 2024, 626, 1025-1033. doi.org/10.1038/s41586-024-07021-y


Maxbridge

Maxbridge is a web-based deterministic graphing program that permits the identification of the maximally bridged ring (or rings) for any molecule using the Chemistry Development Kit (CDK) software library

Maxbridge

Publication:

Marth, C.J., Gallego, G.M., Lee, J.C., Lebold, T.P., Kulyk, S., Kou, K.G.M., Qin, J., Lilien, R. and Sarpong, R., 2015. Network-analysis-guided synthesis of weisaconitine D and liljestrandinine. Nature, 528(7583), pp.493-498. doi.org/10.1038/nature16440


AIMNet2

This package integrates the powerful AIMNet2 neural network potential into your simulation workflows. AIMNet2 provides fast and reliable energy, force, and property calculations for molecules containing a diverse range of elements.

AIMNet2

Publication:

Anstine, D.M., Zubatyuk, R. and Isayev, O., 2025. AIMNet2: a neural network potential to meet your neutral, charged, organic, and elemental-organic needs. Chem. Sci. 2025, 16, 10228-10244. https://doi.org/10.1039/D4SC08572H

Tutorial Material:

AIMNet2 Tutorial

UMA Tutorial


Molcomplex

This package is developed in a collaboration of the Paton and Sarpong groups. It Implements a variety of complementary metrics for molecular complexity and synthetic accessibility.

Molcomplex code

Molcomplex web interface


AiZynthFinder

AiZynthFinder is a free tool for retrosynthetic planning developed at AstraZeneca that C-CAS researchers contributed to. The default algorithm is based on a Monte Carlo tree search that recursively breaks down a molecule to purchasable precursors. The tree search is guided by a policy that suggests possible precursors by utilizing a neural network trained on a library of known reaction templates. This setup is completely customizable as the tool supports multiple search algorithms and expansion policies.

Publications:

Genheden, S., Thakkar, A., Chadimová, V., Reymond, J.L., Engkvist, O., Bjerrum, E., AiZynthFinder: a fast, robust and flexible open-source software for retrosynthetic planning. J. Cheminf. 2020, 12, 70. https://doi.org/10.1186/s13321-020-00472-1

Saigiridharan, L., Hassen, A.K., Lai, H., Torren-Peraire, P., Engkvist, O. and Genheden, S., 2024. AiZynthFinder 4.0: developments based on learnings from 3 years of industrial application. J. Cheminf. 2024 16..57. https://doi.org/10.1186/s13321-024-00860-x

Wiest, O., Bauer, C., Helquist, P., Norrby, P.O. and Genheden, S., 2024. Finding relevant retrosynthetic disconnections for stereocontrolled reactions. J. Chem. Inf. Mod. 2024, 64, 5796-5805. https://doi.org/10.1021/acs.jcim.4c00370

 
Tutorial material:

AiZynthFinder Tutorial