Abstract:Virtual screening (VS) has become a preferred tool to augment high-throughput screening(1) and determine new leads in the drug discovery process. The core of a VS informatics pipeline includes several data mining algorithms that work on huge databases of chemical compounds containing millions of molecular structures and their associated data. Thus, scaling traditional applications such as classification, partitioning, and outlier detection for huge chemical data sets without a significant loss in accuracy is very important. In this paper, we introduce a data mining framework built on top of a recently developed fast approximate nearest-neighbor-finding algorithm(2) called locality-sensitive hashing (LSH) that can be used to mine huge chemical spaces in a scalable fashion using very modest computational resources. The core LSH algorithm hashes chemical descriptors so that points close to each other in the descriptor space are also close to each other in the hashed space. Using this data structure, one can perform approximate nearest-neighbor searches very quickly, in sublinear time. We validate the accuracy and performance of our framework on three real data sets of sizes ranging from 4337 to 249 071 molecules. Results indicate that the identification of nearest neighbors using the LSH algorithm is at least 2 orders of magnitude faster than the traditional k-nearest-neighbor method and is over 94% accurate for most query parameters. Furthermore, when viewed as a data-partitioning procedure, the LSH algorithm lends itself to easy parallelization of nearest-neighbor classification or regression. We also apply our framework to detect outlying (diverse) compounds in a given chemical space this algorithm is extremely rapid in determining whether a compound is located in a sparse region of chemical space or not, and it is quite accurate when compared to results obtained using principal-component-analysis-based heuristics.

Scalable Partitioning and Exploration of Chemical Spaces Using Geometric Hashing

Enhanced Sampling of Chemical Space for High Throughput Screening Applications using Machine Learning

Utilizing Low-Dimensional Molecular Embeddings for Rapid Chemical Similarity Search

Shape-Aware Synthon Search (SASS) for virtual screening of synthon-based chemical spaces

SpaceGrow: efficient shape-based virtual screening of billion-sized combinatorial fragment spaces

Thompson Sampling─An Efficient Method for Searching Ultralarge Synthesis on Demand Databases

Parallel and Distributed Thompson Sampling for Large-scale Accelerated Exploration of Chemical Space

Optimizing substructure search: a novel approach for efficient querying in large chemical databases

Emerging structure-based computational methods to screen the exploding accessible chemical space

Open-Source Approach to GPU-Accelerated Substructure Search

Improving Similarity Search with High-dimensional Locality-sensitive Hashing

A Hadoop-based Massive Molecular Data Storage Solution for Virtual Screening

Experimental Analysis of Locality Sensitive Hashing Techniques for High-Dimensional Approximate Nearest Neighbor Searches

CHEESE: 3D Shape and Electrostatic Virtual Screening in a Vector Space

Customizable Generation of Synthetically Accessible, Local Chemical Subspaces

Density Sensitive Hashing

Virtual Screening of Chemical Space based on Quantum Annealing

Learning-based distributed locality sensitive hashing.

SpaceHASTEN: A structure-based virtual screening tool for non-enumerated virtual chemical libraries

A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening

Efficient Exploration of Chemical Space with Docking and Deep Learning