Abstract:Clustering is a fundamental task in machine learning. One of the most successful and broadly used algorithms is DBSCAN, a density-based clustering algorithm. DBSCAN requires $\epsilon$-nearest neighbor graphs of the input dataset, which are computed with range-search algorithms and spatial data structures like KD-trees. Despite many efforts to design scalable implementations for DBSCAN, existing work is limited to low-dimensional datasets, as constructing $\epsilon$-nearest neighbor graphs is expensive in high-dimensions. In this paper, we modify DBSCAN to enable use of $\kappa$-nearest neighbor graphs of the input dataset. The $\kappa$-nearest neighbor graphs are constructed using approximate algorithms based on randomized projections. Although these algorithms can become inaccurate or expensive in high-dimensions, they possess a much lower memory overhead than constructing $\epsilon$-nearest neighbor graphs. We delineate the conditions under which $k$NN-DBSCAN produces the same clustering as DBSCAN. We also present an efficient parallel implementation of the overall algorithm using OpenMP for shared memory and MPI for distributed memory parallelism. We present results on up to 16 billion points in 20 dimensions, and perform weak and strong scaling studies using synthetic data. Our code is efficient in both low and high dimensions. We can cluster one billion points in 3D in less than one second on 28K cores on the Frontera system at the Texas Advanced Computing Center (TACC). In our largest run, we cluster 65 billion points in 20 dimensions in less than 40 seconds using 114,688 x86 cores on TACC's Frontera system. Also, we compare with a state of the art parallel DBSCAN code; on 20d/4M point dataset, our code is up to 37$\times$ faster.

What problem does this paper attempt to address?

### Problems the Paper Attempts to Solve The paper attempts to address the scalability and efficiency issues of DBSCAN (Density-Based Spatial Clustering of Applications with Noise) on high-dimensional datasets. Specifically, while existing DBSCAN implementations perform well on low-dimensional datasets, they face the following major issues when dealing with high-dimensional data: 1. **Scalability Limitations**: - Constructing the ϵ-nearest neighbor graph (ϵ-NNG) in high-dimensional space is very expensive, with complexity potentially reaching O(𝑑𝑛²), where 𝑑 is the dimensionality of the dataset and 𝑛 is the number of data points. - The ϵ-NNG structure in high-dimensional data is more sensitive to changes in the ϵ parameter, leading to a decline in algorithm performance. 2. **Inefficiency in Parameter Tuning**: - Existing parallel DBSCAN algorithms need to restart the algorithm and perform range queries from scratch when adjusting the input parameters ϵ and minPts, which is very time-consuming on large-scale datasets. - Selecting appropriate parameters typically requires multiple runs of the algorithm, but current methods do not support efficient parameter tuning. To address these issues, the paper proposes a new density clustering algorithm—kNN-DBSCAN. This algorithm uses the k-nearest neighbor graph (k-NNG) to improve efficiency and achieve good scalability on high-dimensional datasets. The specific contributions include: - **Proposing the kNN-DBSCAN Algorithm**: This algorithm improves efficiency and scalability on high-dimensional data by using k-NNG instead of ϵ-NNG. - **Theoretical Proof**: It is proven that kNN-DBSCAN can produce the same results as DBSCAN when using the same input parameters. - **Parallel Implementation**: A hybrid MPI/OpenMP parallel implementation is proposed, using Boruvka's algorithm to construct the minimum spanning tree (MST) and achieving an efficient approximate MST method on distributed memory architectures. - **Experimental Validation**: Experiments demonstrate the superior performance of kNN-DBSCAN on large-scale datasets, especially on high-dimensional datasets, where it is over 37 times faster than existing parallel DBSCAN algorithms. In summary, the paper aims to solve the scalability and parameter tuning efficiency issues of existing DBSCAN on high-dimensional datasets by introducing the kNN-DBSCAN algorithm, providing an efficient and reliable solution for clustering analysis of large-scale data.

KNN-DBSCAN: a DBSCAN in high dimensions

GNN-DBSCAN: A new density-based algorithm using grid and the nearest neighbor

A Novel Density Peaks Clustering Algorithm Based on K Nearest Neighbors with Adaptive Merging Strategy

A Parallel Varied Density-Based Clustering Algorithm with Optimized Data Partition

DBSCAN-KNN-GA: a multi Density-Level Parameter-Free clustering algorithm

A Parallel Adaptive DBSCAN Algorithm Based on k-Dimensional Tree Partition

A Novel DBSCAN Based on Binary Local Sensitive Hashing and Binary-KNN Representation

An efficient and scalable density-based clustering algorithm for datasets with complex structures.

Enabling DBSCAN for Very Large-Scale High-Dimensional Spaces

AMD-DBSCAN: An Adaptive Multi-density DBSCAN for datasets of extremely variable density

Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms

DBSCAN-MS: Distributed Density-Based Clustering in Metric Spaces

A Robust Clustering Algorithm Based on the Identification of Core Points and KNN Kernel Density Estimation

An Efficient Density-based Clustering Algorithm for Higher-Dimensional Data

Efficient index-based KNN join processing for high-dimensional data

A fast DBSCAN algorithm using a bi-directional HNSW index structure for big data

A New K-NN-Centroid-Inspired Density-Based Clustering Algorithm

Exploring Bit-Difference for Approximate KNN Search in High-Dimensional Databases

AutoSCAN: automatic detection of DBSCAN parameters and efficient clustering of data in overlapping density regions

Diagonal Ordering: A New Approach to High-Dimensional KNN Processing.

Parallel Nearest Neighbors in Low Dimensions with Batch Updates