Abstract:In order to improve the accuracy and stability of K-means algorithm and solve the problem of determining the most appropriate number K of clusters and best initial seeds, an improved K-means algorithm based on density Canopy is proposed. Firstly, the density of sample data sets, the average sample distance in clusters and the distance between clusters are calculated, choosing the density maximum sampling point as the first cluster center and removing the density cluster from the data sets. Defining the product of sample density, the reciprocal of the average distance between the samples in the cluster, and the distance between the clusters as weight product, the other initial seeds is determined by the maximum weight product in the remaining data sets until the data sets is empty. The density Canopy is used as the preprocessing procedure of K-means and its result is used as the cluster number and initial clustering center of K-means algorithm. Finally, the new algorithm is tested on some well-known data sets from UCI machine learning repository and on some simulated data sets with different proportions of noise samples. The simulation results show that the improved K-means algorithm based on density Canopy achieves better clustering results and is insensitive to noisy data compared to the traditional K-means algorithm, the Canopy-based K-means algorithm, Semi-supervised K-means++ algorithm and K-means-u* algorithm. The clustering accuracy of the proposed K-means algorithm based on density Canopy is improved by 30.7%, 6.1%, 5.3% and 3.7% on average on UCI data sets, and improved by 44.3%, 3.6%, 9.6% and 8.9% on the simulated data sets with noise signal respectively. With the increase of the noise ratio, the noise immunity of the new algorithm is more obvious, when the noise ratio reached 30%, the accuracy rate is improved 50% and 6% compared to the traditional K-means algorithm and the Canopy-based K-means algorithm.

Improved K-means Clustering Algorithm Based Density and Sample Size

Improved K-means algorithm based on density Canopy

Design and Implementation of an Improved K-Means Clustering Algorithm

An improved k-means algorithm based on density normalization

Improved K average spatial clustering method for nodes of water distribution system

Improved K-means algorithm based on clustering criterion function

An improved K-means algorithm based on multiple feature points

K-means Clustering Algorithm with Improved Initial Center

An Improved K-means Algorithm Based on Mapreduce and Grid

A Novel Effective Distance Measure and a Relevant Algorithm for Optimizing the Initial Cluster Centroids of K-means

An Improved K-Means Clustering Algorithm Based on Spectral Method

An improved K-means algorithm by weighted distance based on the maximum between-cluster variation

An Improved Global K-means Clustering Algorithm

Automatic k-means clustering algorithm based on data sampling

An Improved Density-Based Cluster Analysis Method Combining Genetic Algorithm and Data Sampling for Large-Scale Datasets

A Novel Rough K-means Clustering Algorithm Based on the Weight of Density

A Novel Density Based Clustering Algorithm and Its Parallelization.

Minimum variance and range-based k-means optimization algorithm-Rdk-means algorithm

Improved k-means clustering method for codebook generation

An Improved K-medoids Algorithm Based on Step Increasing and Optimizing Medoids

Improved Density Peak Clustering Algorithm Based on Choosing Strategy Automatically for Cut-off Distance and Cluster Centre