高维分类型数据加权子空间聚类算法

计算机工程与应用 ›› 2014, Vol. 50 ›› Issue (23): 131-135.

• 数据库、数据挖掘、机器学习 • 上一篇下一篇

高维分类型数据加权子空间聚类算法

孙浩军，闪光辉，高玉龙，袁婷，吴云霞

汕头大学工学院，广东汕头 515063

出版日期:2014-12-01 发布日期:2014-12-12

Algorithm for high-dimensional categorical data weighted subspace clustering

SUN Haojun, SHAN Guanghui, GAO Yulong, YUAN Ting, WU Yunxia

College of Engineering, Shantou University, Shantou, Guangdong 515063, China

Online:2014-12-01 Published:2014-12-12

摘要/Abstract

摘要： 子空间聚类是高维数据聚类的一种有效手段，子空间聚类的原理就是在最大限度地保留原始数据信息的同时用尽可能小的子空间对数据聚类。在研究了现有的子空间聚类的基础上，引入了一种新的子空间的搜索方式，它结合簇类大小和信息熵计算子空间维的权重，进一步用子空间的特征向量计算簇类的相似度。该算法采用类似层次聚类中凝聚层次聚类的思想进行聚类，克服了单用信息熵或传统相似度的缺点。通过在Zoo、Votes、Soybean三个典型分类型数据集上进行测试发现：与其他算法相比，该算法不仅提高了聚类精度，而且具有很高的稳定性。

关键词: 高维数据, 聚类, 子空间, 信息熵, 层次聚类

Abstract: Subspace clustering is a kind of effective strategy to high-dimensional data clustering, the principle of subspace clustering is as well as possible keeping original data information, meanwhile as small as possible using subspace to data clustering. Based on the studying of the existing soft subspace clustering, it proposes a new algorithm for subspace searching. The algorithm combines with the size of cluster and information entropy, defines a new subspace dimensional weight distribution mode, and then uses the feature vector of cluster subspace to measure the similarity of two clusters. It uses the idea of agglomerative hierarchical clustering in hierarchical clustering to data clustering, which overcoming the shortcomings of using information entropy or traditional similarity separately. Through the test in the Zoo, Votes, Soybean three typical categorical data set to find out that compared with other algorithms, the proposed algorithm not only can improve the accuracy of clustering, but also has the very high stability.

Key words: high-dimensional data, clustering, subspace, information entropy, hierarchical clustering

孙浩军，闪光辉，高玉龙，袁婷，吴云霞. 高维分类型数据加权子空间聚类算法[J]. 计算机工程与应用, 2014, 50(23): 131-135.

SUN Haojun, SHAN Guanghui, GAO Yulong, YUAN Ting, WU Yunxia. Algorithm for high-dimensional categorical data weighted subspace clustering[J]. Computer Engineering and Applications, 2014, 50(23): 131-135.

[1]	桑江徽，姜海燕. 基于联合分布的多标记迁移学习[J]. 计算机工程与应用, 2021, 57(9): 154-161.
[2]	兰红，黄敏. 融合KNN优化的密度峰值和FCM聚类算法[J]. 计算机工程与应用, 2021, 57(9): 81-88.
[3]	郭晓静，隋昊达. 改进YOLOv3在机场跑道异物目标检测中的应用[J]. 计算机工程与应用, 2021, 57(8): 249-255.
[4]	李莉，纪欣沅，宋嵩. 回环软件缺陷数量预测模型[J]. 计算机工程与应用, 2021, 57(7): 158-163.
[5]	霍光煜，张勇，孙艳丰，尹宝才. 基于语义的档案数据智能分类方法研究[J]. 计算机工程与应用, 2021, 57(6): 247-253.
[6]	杨芳，尹曦，司建辉，刘宏媛，汪雪. 基于侧重点聚类的数学表达式相似度计算方法[J]. 计算机工程与应用, 2021, 57(6): 88-93.
[7]	赵凡，张琳，闻治泉，杨林林，蔺广逢. 一种直接高效的自然场景汉字逼近定位方法[J]. 计算机工程与应用, 2021, 57(6): 159-167.
[8]	彭启慧，宣士斌，高卿. 分布的自动阈值密度峰值聚类算法[J]. 计算机工程与应用, 2021, 57(5): 71-78.
[9]	李勇振，廖湖声. 基于图卷积神经网络的多视角聚类[J]. 计算机工程与应用, 2021, 57(5): 115-122.
[10]	王昌龙，张远东，缪宏，杨煜恒. 双通道卷积神经网络在南瓜病害识别上的应用[J]. 计算机工程与应用, 2021, 57(5): 183-189.
[11]	胡晓敏，王明丰，张首荣，李敏. 用于文本聚类的新型差分进化粒子群算法[J]. 计算机工程与应用, 2021, 57(4): 61-67.
[12]	王鹏，叶学义，王涛，钱丁炜. 双偏差双空间局部方向模式的人脸识别[J]. 计算机工程与应用, 2021, 57(4): 91-99.
[13]	王俊玲，卢新明. 基于语义相关的视频关键帧提取算法[J]. 计算机工程与应用, 2021, 57(4): 192-198.
[14]	王芙银，张德生，张晓. 结合鲸鱼优化算法的自适应密度峰值聚类算法[J]. 计算机工程与应用, 2021, 57(3): 94-102.
[15]	石正宇，陈仁文，黄斌. 基于自归一化神经网络的低分辨率人脸识别[J]. 计算机工程与应用, 2021, 57(3): 137-143.

高维分类型数据加权子空间聚类算法

Algorithm for high-dimensional categorical data weighted subspace clustering

PDF

可视化

摘要/Abstract

引用本文

使用本文

参考文献

相关文章 15

编辑推荐

Metrics