Abstract:Motivation: Effective molecular representation is critical in drug development. The complex nature of molecules demands comprehensive multi-view representations, considering 1D, 2D, and 3D aspects, to capture diverse perspectives. Obtaining representations that encompass these varied structures is crucial for a holistic understanding of molecules in drug-related contexts. Results: In this study, we introduce an innovative multi-view contrastive learning framework for molecular representation, denoted as MolMVC. Initially, we use a Transformer encoder to capture 1D sequence information and a Graph Transformer to encode the intricate 2D and 3D structural details of molecules. Our approach incorporates a novel attention-guided augmentation scheme, leveraging prior knowledge to create positive samples tailored to different molecular data views. To align multi-view molecular positive samples effectively in latent space, we introduce an adaptive multi-view contrastive loss (AMCLoss). In particular, we calculate AMCLoss at various levels within the model to effectively capture the hierarchical nature of the molecular information. Eventually, we pre-train the encoders via minimizing AMCLoss to obtain the molecular representation, which can be used for various down-stream tasks. In our experiments, we evaluate the performance of our MolMVC on multiple tasks, including molecular property prediction (MPP), drug-target binding affinity (DTA) prediction and cancer drug response (CDR) prediction. The results demonstrate that the molecular representation learned by our MolMVC can enhance the predictive accuracy on these tasks and also reduce the computational costs. Furthermore, we showcase MolMVC's efficacy in drug repositioning across a spectrum of drug-related applications. Availability and implementation: The code and pre-trained model are publicly available at https://github.com/Hhhzj-7/MolMVC.

MvMRL: a multi-view molecular representation learning method for molecular property prediction

Self-Supervised Molecular Representation Learning With Topology and Geometry

Multi-Modal Representation Learning for Molecular Property Prediction: Sequence, Graph, Geometry

MolMVC: Enhancing molecular representations for drug-related tasks through multi-view contrastive learning

An improved multi-modal representation-learning model based on fusion networks for property prediction in drug discovery

Learning Multi-view Molecular Representations with Structured and Unstructured Knowledge

Chemical-Reaction-Aware Molecule Representation Learning

MultiModal-Learning for Predicting Molecular Properties: A Framework Based on Image and Graph Structures

Uni-Mol: A Universal 3D Molecular Representation Learning Framework

Multimodal Fusion with Relational Learning for Molecular Property Prediction

A merged molecular representation learning for molecular properties prediction with a web-based service

A focus on molecular representation learning for the prediction of chemical properties

Graph-based Molecular Representation Learning

Graph Multi-Similarity Learning for Molecular Property Prediction

Dual-View Learning Based on Images and Sequences for Molecular Property Prediction

ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction

A Novel Molecular Representation Learning for Molecular Property Prediction with a Multiple SMILES-Based Augmentation

Multi-View Graph Neural Networks for Molecular Property Prediction

Molecular Property Prediction Based on Graph Structure Learning

Molecular Joint Representation Learning via Multi-modal Information