计算机工程与应用 ›› 2016, Vol. 52 ›› Issue (20): 98-102.

• 大数据与云计算 • 上一篇    下一篇

基于聚类的两段式孤立点检测算法

任建华,高立明   

  1. 辽宁工程技术大学 电子与信息工程学院,辽宁 葫芦岛 125105
  • 出版日期:2016-10-15 发布日期:2016-10-14

Two-part outlier detection algorithm based on clustering

REN Jianhua, GAO Liming   

  1. College of Electronics and Information Engineering, Liaoning Technical University, Huludao, Liaoning 125105, China
  • Online:2016-10-15 Published:2016-10-14

摘要: 现有的大多数孤立点检测算法都需要预先设定孤立点个数,并且还缺乏对不均匀数据集的检测能力。针对以上问题,提出了基于聚类的两段式孤立点检测算法,该算法首先用DBSCAN聚类算法产生可疑孤立点集合,然后利用剪枝策略对数据集进行剪枝,并用基于改进距离的孤立点检测算法产生最可能孤立点排序集合,最终由两个集合的交集确定孤立点集合。该算法不必预先设定孤立点个数,具有较高的准确率与检测效率,并且对数据集的分布状况不敏感。数据集上的实验结果表明,该算法能够高效、准确地识别孤立点。

关键词: 孤立点检测, 距离, DBSCAN算法, 剪枝

Abstract: Most of the existing outlier detection algorithms need to preset the number of outliers, and also lack of detection capability of non-uniform data set. In view of the above problems, it puts forward the two-part outlier detection algorithm based on clustering, this algorithm first uses DBSCAN clustering algorithm to produce suspected outlier set, then pruning strategy is used for pruning data set, and the outlier detection algorithm based on improved distance is used to produce the sorting set of the points which most likely to be outliers. Eventually the isolated point set is determined by the intersection of the two sets. The algorithm doesn’t need to preset the number of outliers, with the higher accuracy and detection efficiency, and is not sensitive to the distribution of the data set. The experimental results on data set show that the algorithm can effectively and accurately identify the outliers.

Key words: outlier detection, distance, DBSCAN algorithm, pruning