Abstract:The identification of drug/compound–target interactions (DTIs) constitutes the basis of drug discovery, for which computational predictive approaches have been developed. As a relatively new data-driven paradigm, proteochemometric (PCM) modeling utilizes both protein and compound properties as a pair at the input level and processes them via statistical/machine learning. The representation of input samples (i.e., proteins and their ligands) in the form of quantitative feature vectors is crucial for the extraction of interaction-related properties during the artificial learning and subsequent prediction of DTIs. Lately, the representation learning approach, in which input samples are automatically featurized via training and applying a machine/deep learning model, has been utilized in biomedical sciences. In this study, we performed a comprehensive investigation of different computational approaches/techniques for protein featurization (including both conventional approaches and the novel learned embeddings), data preparation and exploration, machine learning-based modeling, and performance evaluation with the aim of achieving better data representations and more successful learning in DTI prediction. For this, we first constructed realistic and challenging benchmark datasets on small, medium, and large scales to be used as reliable gold standards for specific DTI modeling tasks. We developed and applied a network analysis-based splitting strategy to divide datasets into structurally different training and test folds. Using these datasets together with various featurization methods, we trained and tested DTI prediction models and evaluated their performance from different angles. Our main findings can be summarized under 3 items: (i) random splitting of datasets into train and test folds leads to near-complete data memorization and produce highly over-optimistic results, as a result, should be avoided, (ii) learned protein sequence embeddings work well in DTI prediction and offer high potential, despite interaction-related properties (e.g., structures) of proteins are unused during their self-supervised model training, and (iii) during the learning process, PCM models tend to rely heavily on compound features while partially ignoring protein features, primarily due to the inherent bias in DTI data, indicating the requirement for new and unbiased datasets. We hope this study will aid researchers in designing robust and high-performing data-driven DTI prediction systems that have real-world translational value in drug discovery.

Analysis of protein features and machine learning algorithms for prediction of druggable proteins

Learning the Drug Target-Likeness of A Protein

Support vector machines approach for predicting druggable proteins: recent progress in its exploration and investigation of its usefulness.

Does Drug-Target Have A Likeness?

Prediction of druggable proteins using machine learning and functional enrichment analysis: a focus on cancer-related proteins and RNA-binding proteins

An ensemble method for predicting and designing of druggable proteins.

DrugHybrid_BS: Using Hybrid Feature Combined With Bagging-SVM to Predict Potentially Druggable Proteins

XGB-DrugPred: computational prediction of druggable proteins using eXtreme gradient boosting and optimized features set

Machine learning prediction of oncology drug targets based on protein and network properties

A systematic review of state-of-the-art strategies for machine learning-based protein function prediction

PINNED: identifying characteristics of druggable human proteins using an interpretable neural network

Prediction of Potential Drug Targets Based on Simple Sequence Properties

Machine Learning for Sequence and Structure-Based Protein–Ligand Interaction Prediction

How to approach machine learning-based prediction of drug/compound–target interactions

Testing the predictive power of reverse screening to infer drug targets, with the help of machine learning

In silico prediction of drug-target interaction networks based on drug chemical structure and protein sequences

Prediction Machines: Applied Machine Learning for Therapeutic Protein Design and Development

DrugProtAI: A guide to the future research of investigational target proteins

A deep learning-based method for drug-target interaction prediction based on long short-term memory neural network

DeepConv-DTI: Prediction of drug-target interactions via deep learning with convolution on protein sequences

Exploration of Target Spaces in the Human Genome for Protein and Peptide Drugs