Physics-informed generative model for drug-like molecule conformers

David C. Williams,Neil Inala
DOI: https://doi.org/10.1021/acs.jcim.3c01816
2024-03-15
Abstract:We present a diffusion-based, generative model for conformer generation. Our model is focused on the reproduction of bonded structure and is constructed from the associated terms traditionally found in classical force fields to ensure a physically relevant representation. Techniques in deep learning are used to infer atom typing and geometric parameters from a training set. Conformer sampling is achieved by taking advantage of recent advancements in diffusion-based generation. By training on large, synthetic data sets of diverse, drug-like molecules optimized with the semiempirical GFN2-xTB method, high accuracy is achieved for bonded parameters, exceeding that of conventional, knowledge-based methods. Results are also compared to experimental structures from the Protein Databank (PDB) and Cambridge Structural Database (CSD).
Biomolecules,Machine Learning,Chemical Physics
What problem does this paper attempt to address?
This paper proposes a physics-informed generative model for the generation of drug-like molecular conformations. The model focuses on simulating bond structures and ensures physical consistency using relevant terms from classical force fields. Atom types and geometric parameters are inferred from the training set using deep learning techniques, and conformations are sampled using the latest advances in diffusion-based generation. By training on a large dataset of drug-like molecules optimized with the semi-empirical method GFN2-xTB, the paper achieves high-accuracy prediction of bond parameters, surpassing the accuracy of traditional knowledge-based methods. Additionally, the model's results are compared to experimental structures such as the Protein Data Bank (PDB) and the Cambridge Structural Database (CSD). The main goal of the paper is to establish a conformation generation algorithm that accurately reproduces bond parameters and considers the molecular environment's influence.