Abstract:Abstract Motivation Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs. Results The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and ‘memorize’ the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand. Availability and implementation All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap. Supplementary information Supplementary data are available at Bioinformatics online.

Practically significant method comparison protocols for machine learning in small molecule drug discovery.

Methods to Profile the Macromolecular Targets of Small Compounds.

Machine Learning Small Molecule Properties in Drug Discovery

A call for an industry-led initiative to critically assess machine learning for real-world drug discovery

Implementation of an Automated System Using Machine Learning Models to Accelerate the Process of In Silico Identification of Small Molecules As Drug Candidates

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery

Modern Semiempirical Electronic Structure Methods and Machine Learning Potentials for Drug Discovery: Conformers, Tautomers, and Protonation States

Machine learning for small molecule drug discovery in academia and industry

Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery

Systematic Evaluation of Local and Global Machine Learning Models for the Prediction of ADME Properties

Comparative analysis of machine learning methods in ligand-based virtual screening of large compound libraries.

Towards Evolutionary-based Automated Machine Learning for Small Molecule Pharmacokinetic Prediction

Machine learning in preclinical drug discovery

What are the current challenges for machine learning in drug discovery and repurposing?

Deep Learning Methods for Small Molecule Drug Discovery: A Survey

Impact of Molecular Representations on Deep Learning Model Comparisons in Drug Response Predictions

Validating the validation: reanalyzing a large-scale comparison of deep learning and machine learning models for bioactivity prediction

ALMERIA: Boosting pairwise molecular contrasts with scalable methods

Implementation of The Future of Drug Discovery: QuantumBased Machine Learning Simulation (QMLS)

Modern machine‐learning for binding affinity estimation of protein–ligand complexes: Progress, opportunities, and challenges