Algorithm for high-dimensional categorical data weighted subspace clustering

Abstract

Abstract: Subspace clustering is a kind of effective strategy to high-dimensional data clustering, the principle of subspace clustering is as well as possible keeping original data information, meanwhile as small as possible using subspace to data clustering. Based on the studying of the existing soft subspace clustering, it proposes a new algorithm for subspace searching. The algorithm combines with the size of cluster and information entropy, defines a new subspace dimensional weight distribution mode, and then uses the feature vector of cluster subspace to measure the similarity of two clusters. It uses the idea of agglomerative hierarchical clustering in hierarchical clustering to data clustering, which overcoming the shortcomings of using information entropy or traditional similarity separately. Through the test in the Zoo, Votes, Soybean three typical categorical data set to find out that compared with other algorithms, the proposed algorithm not only can improve the accuracy of clustering, but also has the very high stability.

Key words: high-dimensional data, clustering, subspace, information entropy, hierarchical clustering

摘要： 子空间聚类是高维数据聚类的一种有效手段，子空间聚类的原理就是在最大限度地保留原始数据信息的同时用尽可能小的子空间对数据聚类。在研究了现有的子空间聚类的基础上，引入了一种新的子空间的搜索方式，它结合簇类大小和信息熵计算子空间维的权重，进一步用子空间的特征向量计算簇类的相似度。该算法采用类似层次聚类中凝聚层次聚类的思想进行聚类，克服了单用信息熵或传统相似度的缺点。通过在Zoo、Votes、Soybean三个典型分类型数据集上进行测试发现：与其他算法相比，该算法不仅提高了聚类精度，而且具有很高的稳定性。

关键词: 高维数据, 聚类, 子空间, 信息熵, 层次聚类

SUN Haojun, SHAN Guanghui, GAO Yulong, YUAN Ting, WU Yunxia. Algorithm for high-dimensional categorical data weighted subspace clustering[J]. Computer Engineering and Applications, 2014, 50(23): 131-135.

孙浩军，闪光辉，高玉龙，袁婷，吴云霞. 高维分类型数据加权子空间聚类算法[J]. 计算机工程与应用, 2014, 50(23): 131-135.

[1]	SANG Jianghui, JIANG Haiyan. Multi-label Transfer Learning Algorithm Based on Joint Distribution Alignment [J]. Computer Engineering and Applications, 2021, 57(9): 154-161.
[2]	LAN Hong, HUANG Min. Fusion of KNN Optimized Density Peaks and FCM Clustering Algorithm [J]. Computer Engineering and Applications, 2021, 57(9): 81-88.
[3]	GUO Xiaojing, SUI Haoda. Application of Improved YOLOv3 in Foreign Object Debris Target Detection on Airfield Pavement [J]. Computer Engineering and Applications, 2021, 57(8): 249-255.
[4]	LI Li, JI Xinyuan, SONG Song. Prediction Model for Number of Software Defects in Loop [J]. Computer Engineering and Applications, 2021, 57(7): 158-163.
[5]	HUO Guangyu, ZHANG Yong, SUN Yanfeng, YIN Baocai. Research on Archive Data Intelligent Classification Based on Semantic [J]. Computer Engineering and Applications, 2021, 57(6): 247-253.
[6]	YANG Fang, YIN Xi, SI Jianhui, LIU Hongyuan, WANG Xue. Mathematical Expression Similarity Calculation Method Based on Focus Clustering [J]. Computer Engineering and Applications, 2021, 57(6): 88-93.
[7]	ZHAO Fan, ZHANG Lin, WEN Zhiquan, YANG Linlin, LIN Guangfeng. Direct and Efficient Natural Scene Chinese Character Approaching Spotting Method [J]. Computer Engineering and Applications, 2021, 57(6): 159-167.
[8]	PENG Qihui, XUAN Shibin, GAO Qing. Distribution Automatic Threshold Density Peak Clustering Algorithm [J]. Computer Engineering and Applications, 2021, 57(5): 71-78.
[9]	LI Yongzhen, LIAO Husheng. Multi-view Clustering via Graph Convolutional Neural Network [J]. Computer Engineering and Applications, 2021, 57(5): 115-122.
[10]	WANG Changlong, ZHANG Yuandong, MIAO Hong, YANG Yuheng. Application of Double Channel Convolutional Neural Network in Pumpkin Diseases Identification [J]. Computer Engineering and Applications, 2021, 57(5): 183-189.
[11]	HU Xiaomin, WANG Mingfeng, ZHANG Shourong, LI Min. New Differential Evolution with Particle Swarm Optimization Algorithm for Text Clustering [J]. Computer Engineering and Applications, 2021, 57(4): 61-67.
[12]	WANG Peng, YE Xueyi, WANG Tao, QIAN Dingwei. Face Recognition Based on Double Variation and Double Space Local Directional Pattern [J]. Computer Engineering and Applications, 2021, 57(4): 91-99.
[13]	WANG Junling, LU Xinming. Video Key Frame Extraction Algorithm Based on Semantic Correlation [J]. Computer Engineering and Applications, 2021, 57(4): 192-198.
[14]	WANG Fuyin, ZHANG Desheng, ZHANG Xiao. Adaptive Density Peaks Clustering Algorithm Combining with Whale Optimization Algorithm [J]. Computer Engineering and Applications, 2021, 57(3): 94-102.
[15]	CHEN Junfeng, ZHENG Zhongtuan. Over-Sampling Method on Imbalanced Data Based on WKMeans and SMOTE [J]. Computer Engineering and Applications, 2021, 57(23): 106-112.

Algorithm for high-dimensional categorical data weighted subspace clustering

高维分类型数据加权子空间聚类算法

PDF

Knowledge

Abstract

Cite this article

share this article

References

Related Articles 15

Recommended Articles

Metrics