Abstract:In order to improve the accuracy and stability of K-means algorithm and solve the problem of determining the most appropriate number K of clusters and best initial seeds, an improved K-means algorithm based on density Canopy is proposed. Firstly, the density of sample data sets, the average sample distance in clusters and the distance between clusters are calculated, choosing the density maximum sampling point as the first cluster center and removing the density cluster from the data sets. Defining the product of sample density, the reciprocal of the average distance between the samples in the cluster, and the distance between the clusters as weight product, the other initial seeds is determined by the maximum weight product in the remaining data sets until the data sets is empty. The density Canopy is used as the preprocessing procedure of K-means and its result is used as the cluster number and initial clustering center of K-means algorithm. Finally, the new algorithm is tested on some well-known data sets from UCI machine learning repository and on some simulated data sets with different proportions of noise samples. The simulation results show that the improved K-means algorithm based on density Canopy achieves better clustering results and is insensitive to noisy data compared to the traditional K-means algorithm, the Canopy-based K-means algorithm, Semi-supervised K-means++ algorithm and K-means-u* algorithm. The clustering accuracy of the proposed K-means algorithm based on density Canopy is improved by 30.7%, 6.1%, 5.3% and 3.7% on average on UCI data sets, and improved by 44.3%, 3.6%, 9.6% and 8.9% on the simulated data sets with noise signal respectively. With the increase of the noise ratio, the noise immunity of the new algorithm is more obvious, when the noise ratio reached 30%, the accuracy rate is improved 50% and 6% compared to the traditional K-means algorithm and the Canopy-based K-means algorithm.

Design and Implementation of an Improved K-Means Clustering Algorithm

An Improved K-means Algorithm Based on Multiple Clustering and Density.

Research and Improvement of K-Means Algorithm

An Improved Initial Clustering Center Selection Method for K-Means Algorithm

An Improved Grid-Based K-Means Clustering Algorithm

A Novel Effective Distance Measure and a Relevant Algorithm for Optimizing the Initial Cluster Centroids of K-means

An improved dynamic K-means clustering algorithm

Improved Initial Cluster Center Selection in K-Means Clustering

An Improved K-Means Algorithm Based on Kurtosis Test

Improved k-means algorithm with meliorated initial centers

K-means Clustering Algorithm Based on Initial Clustering Centre Selection and Points Division

Improved K-means Clustering Algorithm Based Density and Sample Size

An Improved Global K-means Clustering Algorithm

An Improved Clustering Algorithm Based on Cluster Weight Coefficient

K-means Clustering Algorithm with Improved Initial Center

Improved K-means algorithm based on density Canopy

K*-Means: an Effective and Efficient K-Means Clustering Algorithm

An Improved Method Based On The Density And K-Means Nearest Neighbor Text Clustering Algorithm

An Improved Clustering Method Based on Data Field

New K-Means Clustering Center Select Algorithm

An Improved Rough K-Means Algorithm with Weighted Distance Measure