Abstract:We explore transfer learning models from a pre-trained graph convoluntional neural network representation of molecules, obtained from SchNet, 1 to predict 13 C-NMR, pKa, and logS sol- ubility. SchNet learns a graph representation of a molecule by associating each atom with an “embedding vector” and interacts the atom-embeddings with each other by leveraging graph- convolutional filters on their interatomic distances. We pre-trained SchNet on molecular energy and demonstrate that the pre-trained atomistic embeddings can then be used as a transferable representation for a wide array of properties. On the one hand, for atomic properties such as micro-pK1 and 13 C-NMR, we investigate two models, one linear and one neural net, that inputs pre-trained atom-embeddings of a particular atom (e.g. carbon) and predicts a local property (e.g. 13 C-NMR). On the other hand, for molecular properties such as solubility, a size-extensive graph model is built using the embeddings of all atoms in the molecule as input. For all cases, qualitatively correct predictions are made with relatively little training data (< 1000 training points), showcasing the ease with which pre-trained embeddings pick up on important chemical patterns. The proposed models successfully capture well-understood trends of pK1 and solu- bility. This study advances our understanding of current neural net graph representations and their capacity for transfer learning applications in chemistry.

What problem does this paper attempt to address?

This paper discusses how to use transfer learning to predict the 13C-NMR, pKa, and solubility of molecules. The study is based on a pre-trained graph convolutional neural network (GCNN) called SchNet, which associates each atom with an "embedding vector" and utilizes graph convolution filters on the distances between atoms to interact these atom embedding vectors and learn the graph representation of molecules. The authors first pre-train SchNet on molecular energies, and then demonstrate that these pre-trained atom embedding vectors can be used as transferable representations for predicting various properties. For atomic properties such as 13C-NMR and pKa, the researchers use linear and neural network models with input-specific atom (e.g., carbon) pre-trained atom embedding vectors to predict local properties. For molecular properties like solubility, they construct a scale-expanding graph model that takes all atom embedding vectors as input. These models can make qualitatively correct predictions with relatively limited training data (less than 1000 training points), showcasing the ability of pre-trained embedding vectors to capture important chemical patterns quickly. The paper highlights the advantages of pre-trained embedding vectors in handling small datasets, which is an improvement over challenges faced by traditional machine learning methods. Additionally, the study explores how to perform transfer learning using pre-trained graph representations to predict a wider range of chemical properties, aiming to achieve a universal representation that can handle almost any property, similar to deriving all molecular properties from a molecular wave function. Experimental results show that linear models achieve satisfactory accuracy in predicting pKa and 13C-NMR, although non-linear models provide slight improvements, the relationship is primarily linear. For solubility, the study proposes a scale-expanding neural network model that predicts contributions from the atomic level, enabling predictions of solubility for the entire molecule. In summary, the paper aims to address the effective prediction of chemical properties, particularly when limited data is available, using pre-trained graph neural network models through transfer learning to improve predictive performance.

Transfer Learning Graph Representations of Molecules for pKa, 13C-NMR, and Solubility

Transfer Learning Graph Representations of Molecules for pKa, 13C-NMR, and Solubility

MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction

Leveraging Graph Neural Networks for Enhanced Prediction of Molecular Solubility via Transfer Learning

Transfer learning with graph neural networks for improved molecular property prediction in the multi-fidelity setting

Accurate Predictions of Molecular Properties of Proteins via Graph Neural Networks and Transfer Learning

An Interpretable Graph Representation Learning Model for Accurate Predictions of Drugs Aqueous Solubility

Accurate predictions of aqueous solubility of drug molecules via the multilevel graph convolutional network (MGCN) and SchNet architectures

Transfer learning based on atomic feature extraction for the prediction of experimental ¹³C chemical shifts

Transfer Learning for Molecular Property Predictions from Small Data Sets

Convolutional Embedding of Attributed Molecular Graphs for Physical Property Prediction

Graph Neural Tree: A novel and interpretable deep learning-based framework for accurate molecular property predictions

Neural Mulliken Analysis: Molecular Graphs from Density Matrices for QSPR on Raw Quantum-Chemical Data

Enhancing Model Learning and Interpretation Using Multiple Molecular Graph Representations for Compound Property and Activity Prediction

Few-shot learning via graph embeddings with convolutional networks for low-data molecular property prediction

Describe Molecules by a Heterogeneous Graph Neural Network with Transformer-like Attention for Supervised Property Predictions

SolvBERT for solvation free energy and solubility prediction: a demonstration of an NLP model for predicting the properties of molecular complexes

Transfer learning for atomistic simulations using GNNs and kernel mean embeddings

Leveraging 2D molecular graph pretraining for improved 3D conformer generation with graph neural networks

Predicting Aqueous Solubility of Organic Molecules Using Deep Learning Models with Varied Molecular Representations

Nucleophilicity Prediction Using Graph Neural Networks