Improved K-means algorithm based on global center and nonuniqueness high-density points

Computer Engineering and Applications ›› 2016, Vol. 52 ›› Issue (1): 48-54.

Previous Articles Next Articles

Improved K-means algorithm based on global center and nonuniqueness high-density points

HE Yunbin1, LIU Xuejiao1, WANG Zhiqiang2, WAN Jing1, LI Song1

1.School of Computer Science and Technology, Harbin University of Science and Technology, Harbin 150080, China
2.Department of Computer and Information Engineering, Heilongjiang University of Finance and Economics, Harbin 150025, China

Online:2016-01-01 Published:2015-12-30

基于全局中心的高密度不唯一的K-means算法研究

何云斌1，刘雪娇1，王知强2，万静1，李松1

1.哈尔滨理工大学计算机科学与技术学院，哈尔滨 150080
2.黑龙江财经学院计算机与信息工程系，哈尔滨 150025

Abstract

Abstract: The traditional K-means algorithm is sensitive to the selection of the initial clustering center, moreover the clustering number-k can not be confirmed beforehand. Then it is not conductive to the stability of the clustering. Considering this defection, a new improved algorithm—NDK-means（Nonuniqueness high-Density K-means） is proposed which is based on the global clustering center. In this algorithm, the highest-density points are not unique. To determine the radius, the standard deviation is applied to the new algorithm, then the initial clustering centers can be selected from a set of the high-density areas of points. When the highest-density points are not unique, the proposed new algorithm selects a set of initial clustering centers which are the furthest distance from the global clustering center. Besides, the new algorithm can determine the optimal number of cluster by selecting the BWP validity index. Finally, the experimental and analysis results show the new proposed method outperforms than the traditional K-means algorithm in terms of the validity and stability.

Key words: K-means algorithm, initial clustering center, number of clusters, density-based

摘要： 传统的K-means算法敏感于初始中心点的选取，并且无法事先确定准确的聚类数目[k]，不利于聚类结果的稳定性。针对传统K-means算法的以上不足，提出了基于全局中心的高密度不唯一的新方法——NDK-means，该方法通过标准差确定有效密度半径，并从高密度区域中选取具有代表性的样本点作为初始聚类中心。此外算法针对最高密度点不唯一的情况进行特别分析，选取距离全局中心最远的点集作为最优的初始中心点集合。在NDK-means算法基础上结合有效性指标BWP对聚类结果进行分析，从而解决了最佳有效聚类数目无法事先确定的不足。理论研究与实验结果表明所提方法的聚类结果具有更好的稳定性和可行性。

关键词: K-means算法, 初始中心, 聚类数, 基于密度

HE Yunbin1, LIU Xuejiao1, WANG Zhiqiang2, WAN Jing1, LI Song1. Improved K-means algorithm based on global center and nonuniqueness high-density points[J]. Computer Engineering and Applications, 2016, 52(1): 48-54.

何云斌1，刘雪娇1，王知强2，万静1，李松1. 基于全局中心的高密度不唯一的K-means算法研究[J]. 计算机工程与应用, 2016, 52(1): 48-54.

[1]	PAN Chengsheng, ZHANG Bin, LYU Yana, DU Xiuli, QIU Shaoming. K-Means Text Clustering Based on Improved Gray Wolf Optimization Algorithm [J]. Computer Engineering and Applications, 2021, 57(1): 188-193.
[2]	WANG Zilong, LI Jin, SONG Yafei. Improved K-means Algorithm Based on Distance and Weight [J]. Computer Engineering and Applications, 2020, 56(23): 87-94.
[3]	ZHANG Zhen, LI Haofang, LI Mengzhou. Research on YOLO Algorithm in Abnormal Security Images [J]. Computer Engineering and Applications, 2020, 56(21): 187-193.
[4]	WANG Liang, YE Jimin. Hybrid Algorithm of DBSCAN and Improved SMOTE for Oversampling [J]. Computer Engineering and Applications, 2020, 56(18): 111-118.
[5]	GUO Yongkun, ZHANG Xinyou, LIU Liping, DING Liang, NIU Xiaolu. K-means Clustering Algorithm of Optimizing Initial Clustering Center [J]. Computer Engineering and Applications, 2020, 56(15): 172-178.
[6]	LI Feng, LI Mingxiang, ZHANG Yujing. Partial Iterative Fast K-means Clustering Algorithm [J]. Computer Engineering and Applications, 2020, 56(13): 63-71.
[7]	WANG Jianren, MA Xin, DUAN Ganglong. Improved K-means Clustering k-Value Selection Algorithm [J]. Computer Engineering and Applications, 2019, 55(8): 27-33.
[8]	HU Jian1, ZHU Haiwan2, MAO Yimin2. DBSCAN Clustering Algorithm Based on Adaptive Bee Colony Optimization [J]. Computer Engineering and Applications, 2019, 55(14): 105-114.
[9]	CHEN Qinghu, ZHOU Xiaodan, YAN Yuchen. Recognition of print file based on character image segmentation [J]. Computer Engineering and Applications, 2018, 54(7): 170-175.
[10]	ZHANG Wenyuan, TAN Guoxin, ZHU Xiangzhou. Application of stay points spatial clustering in hot scenic spots analysis [J]. Computer Engineering and Applications, 2018, 54(4): 263-270.
[11]	ZHOU Benjin, TAO Yizheng, JI Bin, XIE Yonghui. Optimizing k-means initial clustering centers by minimizing sum of squared error [J]. Computer Engineering and Applications, 2018, 54(15): 48-52.
[12]	WANG Binyu1, LIU Wenfen2, HU Xuexian1, WEI Jianghong1. Research on text clustering for selecting initial cluster center based on Cosine distance [J]. Computer Engineering and Applications, 2018, 54(10): 11-18.
[13]	WANG Zhaofeng, SHAN Ganlin . k-means based method for dynamically selecting DBSCAN algorithm parameters [J]. Computer Engineering and Applications, 2017, 53(3): 80-86.
[14]	BAI Shuren1，2, CHEN Long2. Particle clustering algorithm with adaptive K values [J]. Computer Engineering and Applications, 2017, 53(16): 116-120.
[15]	CHEN Leilei, GE Hongwei, YANG Jinlong, YUAN Yunhao. Manifold structure based multi-exemplar affinity propagation [J]. Computer Engineering and Applications, 2016, 52(6): 67-73.

Improved K-means algorithm based on global center and nonuniqueness high-density points

基于全局中心的高密度不唯一的K-means算法研究

PDF

Knowledge

Abstract

Cite this article

share this article

References

Related Articles 15

Recommended Articles

Metrics