Abstract:Motivation: Virtual screening of molecular compound libraries is a potentially powerful and inexpensive method for the discovery of novel lead compounds for drug development. The major weakness of virtual screening-the inability to consistently identify true positives (leads)-is likely due to our incomplete understanding of the chemistry involved in ligand binding and the subsequently imprecise scoring algorithms. It has been demonstrated that combining multiple scoring functions (consensus scoring) improves the enrichment of true positives. Previous efforts at consensus scoring have largely focused on empirical results, but they have yet to provide a theoretical analysis that gives insight into real features of combinations and data fusion for virtual screening. Results: We demonstrate that combining multiple scoring functions improves the enrichment of true positives only if (a) each of the individual scoring functions has relatively high performance and (b) the individual scoring functions are distinctive. Notably, these two prediction variables are previously established criteria for the performance of data fusion approaches using either rank or score combinations. This work, thus, establishes a potential theoretical basis for the probable success of data fusion approaches to improve yields in in silico screening experiments. Furthermore, it is similarly established that the second criterion (b) can, in at least some cases, be functionally defined as the area between the rank versus score plots generated by the two (or more) algorithms. Because rank-score plots are independent of the performance of the individual scoring function, this establishes a second theoretically defined approach to determining the likely success of combining data from different predictive algorithms. This approach is, thus, useful in practical settings in the virtual screening process when the performance of at least two individual scoring functions (such as in criterion a) can be estimated as having a high likelihood of having high performance, even if no training sets are available. We provide initial validation of this theoretical approach using data from five scoring systems with two evolutionary docking algorithms on four targets, thymidine kinase, human dihydrofolate reductase, and estrogen receptors of antagonists and agonists. Our procedure is computationally efficient, able to adapt to different situations, and scalable to a large number of compounds as well as to a greater number of combinations. Results of the experiment show a fairly significant improvement (vs single algorithms) in several measures of scoring quality, specifically "goodness-of-hit" scores, false positive rates, and "enrichment". This approach (available online at http://gemdock.life. nctu.edu.tw/dock/download.php) has practical utility for cases where the basic tools are known or believed to be generally applicable, but where specific training sets are absent.

Accuracy or novelty: what can we gain from target-specific machine-learning-based scoring functions in virtual screening?

Beware of the Generic Machine Learning-Based Scoring Functions in Structure-Based Virtual Screening.

A Case-Based Meta-Learning Algorithm Boosts the Performance of Structure-Based Virtual Screening.

Can Machine Learning Consistently Improve the Scoring Power of Classical Scoring Functions? Insights into the Role of Machine Learning in Scoring Functions.

Recent progress on the prospective application of machine learning to structure-based virtual screening

Assessment of the Generalization Abilities of Machine-Learning Scoring Functions for Structure-Based Virtual Screening

A Small Step Toward Generalizability: Training a Machine Learning Scoring Function for Structure-Based Virtual Screening

TB-IECS: an accurate machine learning-based scoring function for virtual screening

Featurization strategies for protein–ligand interactions and their applications in scoring function development

Machine‐learning scoring functions for structure‐based drug lead optimization

Improving Structure-Based Virtual Screening Performance Via Learning from Scoring Function Components

Machine Learning Scoring Functions for Drug Discoveries from Experimental and Computer-Generated Protein-Ligand Structures: Towards Per-Target Scoring Functions

From machine learning to deep learning: Advances in scoring functions for protein–ligand docking

Data-augmented machine learning scoring functions for virtual screening of YTHDF1 m6A reader protein

Consensus scoring criteria for improving enrichment in virtual screening

Can molecular dynamics simulations improve predictions of protein-ligand binding affinity with machine learning?

Docking Score ML: Target-Specific Machine Learning Models Improving Docking-Based Virtual Screening in 155 Targets

A Generalized Protein-Ligand Scoring Framework with Balanced Scoring, Docking, Ranking and Screening Powers.

A comprehensive comparative assessment of 3D molecular similarity tools in ligand-based virtual screening

True Accuracy of Fast Scoring Functions to Predict High-Throughput Screening Data from Docking Poses: The Simpler the Better

Computational representations of protein–ligand interfaces for structure-based virtual screening