Resources

  • AcidAmine Descriptor Predict

    This repository contains the necessary code for work done in the paper 'Rapid Prediction of Conformationally-Dependent DFT-Level Descriptors using Graph Neural Networks for Carboxylic Acids and Alkyl Amines'

    View details
  • AgentDrug

    Molecular editing—modifying a given molecule to improve desired properties—is a fundamental task in drug discovery. While LLMs hold the potential to solve this task using natural language to drive the editing, straightforward prompting ach…

    View details
  • AIMNet2

    This package integrates the powerful AIMNet2 neural network potential into your simulation workflows. AIMNet2 provides fast and reliable energy, force, and property calculations for molecules containing a diverse range of elements.

    View details
  • AiZynthFinder

    AiZynthFinder is a free tool for retrosynthetic planning developed at AstraZeneca that C-CAS researchers contributed to.

    View details
  • ALFABET

    This library contains the trained graph neural network model for the prediction of homolytic bond dissociation energies (BDEs) of organic molecules with C, H, N, and O atoms. This package offers a command-line interface to the web-based mo…

    View details
  • Alkyl Amines

    Alkyl Amines Feature Library

    View details
  • Anilines

    Anilines Feature Library

    View details
  • AQME

    AQME is an ensemble of automated QM workflows, including: 1) RDKit- and CREST-based conformer generator and ready-to-submit QM input files starting from individual files or databases, 2) post-processing of QM output files to fix extra imag…

    View details
  • Aryl Bromides

    Aryl Bromides Feature Library

    View details
  • Auto-QChem

    Auto-QChem is an automatic, high-throughput and end-to-end DFT calculation workflow that computes chemical descriptors for organic molecules. Tailored toward users without extensive programming experience, Auto-QChem has facilitated more t…

    View details
  • Bandit-Optimization

    This code uses reinforcement learning, specifically the multi-armed bandit approach, for reaction optimization. It demonstrates data-efficient learning at high accuracies and has unique functionalities.

    View details
  • BDE-db

    A database of 290,664 bond dissociation energies for small molecules (10 or fewer heavy atoms) consisting of C, H, O, or N atoms.

    View details
  • Bisphosphine Conformer Selection

    A foundational consideration in the development of computationally derived molecular feature libraries is the generation and selection of conformers. It has been shown that several feature values have a degree of conformer depencency – whi…

    View details
  • Bisphosphine Data

    Available code for the investigation into conformational dependace of features for Pd[allyl] bisphosphine complexes.

    View details
  • BUNNY: An N, N-Bidentate Nitrogen Ligand Descriptor Library

    A density functional theory-based descriptor library of approximately 1100 N,N-bidentate ligands designed to support modeling tasks for Ni-catalyzed cross-coupling.

    View details
  • Carboxylic Acids

    Carboxylic Acids Feature Library

    View details
  • CASCADE

    CASCADE stands for ChemicAl Shift CAlculation with DEep learning. It is a stereochemistry-aware online calculator for NMR chemical shifts using a graph network approach developed at Colorado State University.

    View details
  • ChemOrch

    ChemOrch is a framework that synthesizes chemically grounded instruction-response pairs through a two-stage process: task-controlled instruction generation and tool-aware response construction. ChemOrch enables controllable diversity and l…

    View details
  • Cyanoarenes

    Cyanoarenes Feature Library

    View details
  • DA DataExtraction

    A series of Jupyter notebooks and instructions for extracting a dataset of Diels–Alder reactions.

    View details
  • DBStep

    DBStep is a python package for obtaining DFT-Based Steric Parameters from 3-dimensional chemical structures. It can parse the outputs from most computational chemistry programs and other common molecular structure file formats.

    View details
  • DESP: Double-Ended Synthesis Planning with Goal-Constrained Bidirectional Search

    This repository contains code for DESP (Double-Ended Synthesis Planning), which applies goal-constrained bidirectional search to computer-aided synthesis planning.

    View details
  • Desulfonylative Fluorination

    Modeling scripts and data associated with the Sigman-Sanford lab collaboration on desulfonylative fluorination of heteroaromatics. DOI: pending

    View details
  • DoMiNO

    DoMiNO is a multi-scale framework that decomposes MD dynamics into several temporal resolutions, each governed by a neural graph ordinary differential equation (GraphODE) and is adaptively fused for final predictions.

    View details
  • EDBO+

    EDBO+ is a multi-objective reaction Bayesian optimization platform that builds on the previously published Bayesian optimizers EDBO. The web-based application incorporates features such as condition modification on the fly and data visuali…

    View details
  • GoodVibes

    GoodVibes is a Python program to compute thermochemical data from one or a series of electronic structure calculations. I

    View details
  • Hands-On Data Science for Chemists

    This book serves as a practical introduction to the integration of data science and chemistry. Designed specifically for chemists, it bridges the gap between these fields, offering step-by-step tutorials and real-world applications to tack…

    View details
  • Higher-Level Strategies for Computer-Aided Retrosynthesis

    Retrosynthesis is a core technique in organic chemistry that simplifies target molecules into more readily available components. Computer-aided synthesis planning (CASP) automates this process by recursively proposing immediate precursors …

    View details
  • HT TSs Opt

    This repository contains scripts for high-throughput generation, optimization, and featurization of TSs and catalytic cycle intermediates, along with scripts for MLR and active learning modeling and Excel spreadsheets with input data.

    View details
  • Kraken

    Kraken is a discovery platform covering monodentate organophosphorus(III) ligands providing comprehensive physicochemical descriptors based on representative conformer ensembles.

    View details
  • LabSafety Bench

    LabSafety Bench is a comprehensive evaluation framework designed to rigorously assess the trustworthiness of large language models in laboratory settings. The benchmark includes two main evaluation components:

    View details
  • LLM Extraction Chem

    In the realm of chemistry, literature texts elucidating chemical reactions are crucial for tasks such as yield prediction, reaction prediction, and reaction condition recommendation. However, extracting structured data from these texts is …

    View details
  • MARCEL

    MARCEL is a PyTorch-based benchmark library that evaluates the potential of machine learning on conformer ensembles across a diverse set of molecules, datasets, and models.

    View details
  • Maxbridge

    Maxbridge is a web-based deterministic graphing program that permits the identification of the maximally bridged ring (or rings) for any molecule using the Chemistry Development Kit (CDK) software library

    View details
  • Molcomplex

    This package is developed in a collaboration of the Paton and Sarpong groups. It Implements a variety of complementary metrics for molecular complexity and synthetic accessibility.

    View details
  • MolPuzzle

    MolPuzzle is a benchmark comprising 234 instances of structure elucidation, which feature over 18,000 QA samples presented in a sequential puzzle-solving process, involving three interlinked subtasks: molecule understanding, spectrum inter…

    View details
  • Multi-Linear Regression

    Multilinear regressions (MLR) is a widely used tool in predictive chemistry and the elucidation of mechanism. C-CAS researchers developed a variety of tools for conducting and visualizing MLR analyses, the rapid generation of chemically re…

    View details
  • Multi-Threshold Analysis

    This repository contains a workflow for classification of chemical compounds as active or inactive based on experimental outputs and a set of previously computed descriptors. Classification is performed via sci-kit learn's DecisionTreeClas…

    View details
  • One Step Retro Failure Mode

    This repository contains code and scripts for analyzing and quantifying failure modes of one-step retrosynthesis models, as described in the publication below. Please refer to the paper for detailed methodology and results.

    View details
  • Open Reaction Database

    The Open Reaction Database (ORD) is an open-access schema and infrastructure for structuring and sharing organic reaction data, including a centralized data repository. The ORD schema supports conventional and emerging technologies, from b…

    View details
  • PATRO

    This repository contains the code for Pathway-Aware Template-Based Retrosynthesis (DOI: 10.1021/acs.jcim.6c01458). This single-step model augments the template relevance model from ASKCOS to consider the reaction pathway history when makin…

    View details
  • PericyclicTL

    Data generation notebooks for pericyclic datasets. These notebooks are for regenerating reaction data used in an upcoming publication.

    View details
  • Proto-Yield

    Proto-Yield is an encoder-agnostic prototype network that models reactions as occurring in one of three yield regimes: high, medium, or low. Without access to full reaction processes, Proto-Yield learns to infer latent regimes and their as…

    View details
  • Q2MM

    Q2MM stands for quantum (mechanics) to molecular mechanics or quantum guided molecular mechanics, depending on what you prefer. Q2MM is open source software for force field optimization.

    View details
  • Quinones

    Quinones Feature Library

    View details
  • ReactionTeam

    ReactionTeam, is composed of specialized expert models, each trained to capture a distinct type of electron redistribution pattern in reaction, and a ranking expert that evaluates and orders the generated predictions

    View details
  • REyes (Reciprocal Eyes)

    This repository contains a series of scripts designed to process diffraction data, generate heatmaps, identify key targets, and manage navigation files for SerialEM. The scripts are structured to be executed in the order outlined below, en…

    View details
  • RiskLab

    When multiple LLM agents interact — negotiating prices, relaying information, or making collective decisions — new risks emerge from the interaction itself, not from any single agent's failure. Agents may silently collude on prices, confor…

    View details
  • Rxnpredict

    Predicting reaction performance using machine learning

    View details
  • SMART Molecular Descriptors

    An open-source Python package for generation of SMART probe pockets and calculation of molecular descriptors.

    View details
  • SPiCE: Symmetry-Preserving Conformer Ensemble Networks for Molecular Representation Learning

    SPiCE learns molecular properties from conformer ensembles while preserving joint equivariance to geometric transformations of individual conformers and permutations of the ensemble.

    View details
  • Sulfides

    Sulfides Feature Library

    View details
  • Sulfonates

    Sulfonates Feature Libraries

    View details
  • Sulfonimidamides

    Given the lack of commercially available racemic sulfonimidamides, chemists generated a list totaling 117 diverse, synthetically-feasible sulfonimidamides

    View details
  • Sulfonyl Fluorides

    Sulfonyl Fluorides Feature Library

    View details
  • Threshold

    Python tool to assess data for single-parameter thresholds

    View details
  • TrustGen

    TrustGen is the first dynamic benchmarking platform designed to evaluate trustworthiness across multiple dimensions and model types, including text-to-image, large language, and vision-language models. TrustGen leverages modular components…

    View details
  • Unactivated Primary Aryl Bromides

    The initial library of primary alkyl bromides was selected from the Auto-QChem database developed by the Doyle Lab. After excluding alpha-carbonyl, benzylic, allylic, propargylic bromides, and alkyl bromides containing iodide, a curated se…

    View details
  • wSterimol

    wSterimol is an automated computational workflow which computes multidimensional Sterimol parameters.

    View details
  • yield-rxn

    Code for the paper: Graph Neural Networks for Predicting Chemical Reaction Performance

    View details