计算机工程与应用 ›› 2013, Vol. 49 ›› Issue (20): 118-121.

• 数据库、数据挖掘、机器学习 • 上一篇    下一篇

不均衡数据集上文本分类方法研究

谢娜娜,房  斌,吴  磊   

  1. 重庆大学 计算机学院,重庆 400030
  • 出版日期:2013-10-15 发布日期:2013-10-30

Study of text categorization on imbalanced data

XIE Nana, FANG Bin, WU Lei   

  1. College of Computer Science, Chongqing University, Chongqing 400030, China
  • Online:2013-10-15 Published:2013-10-30

摘要: 文本分类中数据集的不均衡问题是一个在实际应用中普遍存在的问题。从特征选择优化和分类器性能提升两方面出发,提出了一种组合的不均衡数据集文本分类方法。在特征选择方面,综合考虑特征项与类别的正负相关特性及类别区分强度对传统CHI统计特征选择方法予以改进。在数据层上,采用数据重取样方法对不均衡训练语料的不平衡性过滤减少其对分类性能的影响。实验结果表明该方法对不均衡数据集上文本可达到较好分类效果。

关键词: 特征选择, CHI统计, 文本分类, 不均衡数据集, 重取样

Abstract: Class imbalance problems are often encountered in real application of automatic text classifications. From the view of the optimistic feature selection methods and the improvement of classifiers, a new text classification method on imbalanced data set is proposed. The positive and negative correlation between items and categorizations are combined with the strength of class information in the aspect of the feature selection scheme. Then on the data layer, the imbalanced characters of the training corpus are filtered by data resampling methods in order to reduce the effect on the classification. Experimental results show that the?new approach can achieve better performance.

Key words: feature selection, CHI statistical approach, text categorization, imbalanced data;resampling