Learning Chemical Reaction Representation with Reactant-Product Alignment

Kaipeng Zeng,Xianbin Liu,Yu Zhang,Xiaokang Yang,Yaohui Jin,Yanyan Xu

2024-11-27

Abstract:Organic synthesis stands as a cornerstone of chemical industry. The development of robust machine learning models to support tasks associated with organic reactions is of significant interest. However, current methods rely on hand-crafted features or direct adaptations of model architectures from other domains, which lacks feasibility as data scales increase or overlook the rich chemical information inherent in reactions. To address these issues, this paper introduces {\modelname}, a novel chemical reaction representation learning model tailored for a variety of organic-reaction-related tasks. By integrating atomic correspondence between reactants and products, our model discerns the molecular transformations that occur during the reaction, thereby enhancing the comprehension of the reaction mechanism. We have designed an adapter structure to incorporate reaction conditions into the chemical reaction representation, allowing the model to handle diverse reaction conditions and adapt to various datasets and downstream tasks, e.g., reaction performance prediction. Additionally, we introduce a reaction-center aware attention mechanism that enables the model to concentrate on key functional groups, thereby generating potent representations for chemical reactions. Our model has been evaluated on a range of downstream tasks, including reaction condition prediction, reaction yield prediction, and reaction selectivity prediction. Experimental results indicate that our model markedly outperforms existing chemical reaction representation learning architectures across all tasks. Notably, our model significantly outperforms all the baselines with up to 25\% (top-1) and 16\% (top-10) increased accuracy over the strongest baseline on USPTO\_CONDITION dataset for reaction condition prediction. We plan to open-source the code contingent upon the acceptance of the paper.

Machine Learning,Artificial Intelligence

What problem does this paper attempt to address?

The problem that this paper attempts to solve is how to effectively represent chemical reactions in organic chemical reactions to support tasks related to organic reactions, such as reaction condition prediction, reaction yield prediction, and reaction selectivity prediction. Current methods either rely on hand - designed features or directly borrow model architectures from other fields. These methods lack feasibility when the data scale expands or ignore the rich chemical information in the reactions. To overcome these problems, the paper proposes RAlign, a new chemical reaction representation learning model, which improves the understanding of reaction mechanisms by integrating the atomic correspondence between reactants and products. In addition, RAlign also designs an adapter structure to integrate reaction conditions, and a reaction - center - aware attention mechanism, enabling the model to focus on key functional groups, thereby generating powerful chemical reaction representations.

Learning Chemical Reaction Representation with Reactant-Product Alignment

Exploring Chemical Reaction Space with Machine Learning Models: Representation and Feature Perspective

Prediction of Organic Reaction Outcomes Using Machine Learning

Self-Supervised Contrastive Molecular Representation Learning with a Chemical Synthesis Knowledge Graph

Bridging the Gap between Chemical Reaction Pretraining and Conditional Molecule Generation with a Unified Model

ReaKE: Contrastive Molecular Representation Learning with Chemical Synthetic Knowledge Graph

Unified Deep Learning Model for Multitask Reaction Predictions with Explanation

Developing Deep Learning-Based Large-Scale Organic Reaction Classification Model Via Sigma-Profiles

Beyond the Typical: Modeling Rare Plausible Patterns in Chemical Reactions by Leveraging Sequential Mixture-of-Experts

Bridging Chemical Knowledge and Machine Learning for Performance Prediction of Organic Synthesis.

Generic Interpretable Reaction Condition Predictions with Open Reaction Condition Datasets and Unsupervised Learning of Reaction Center

Enhancing chemical synthesis: a two-stage deep neural network for predicting feasible reaction conditions

Root-aligned SMILES: a Tight Representation for Chemical Reaction Prediction

Chemical-Reaction-Aware Molecule Representation Learning

Reaction performance prediction with an extrapolative and interpretable graph model based on chemical knowledge

ReLM: Leveraging Language Models for Enhanced Chemical Reaction Prediction

Beyond Major Product Prediction: Reproducing Reaction Mechanisms with Machine Learning Models Trained on a Large-Scale Mechanistic Dataset

A generalized-template-based graph neural network for accurate organic reactivity prediction

Machine-Learning-Guided Discovery of Electrochemical Reactions

Using Machine Learning To Predict Suitable Conditions for Organic Reactions

Reacon: a template- and cluster-based framework for reaction condition prediction