Empirical study of sentiment classification for Chinese microblog based on machine learning

Computer Engineering and Applications ›› 2012, Vol. 48 ›› Issue (1): 1-4.

• 博士论坛 • Previous Articles Next Articles

Empirical study of sentiment classification for Chinese microblog based on machine learning

LIU Zhiming, LIU Lu

School of Economics and Management, Beihang University, Beijing 100191, China

Received:1900-01-01 Revised:1900-01-01 Online:2012-01-01 Published:2012-01-01

基于机器学习的中文微博情感分类实证研究

刘志明，刘鲁

北京航空航天大学经济管理学院，北京 100191

Abstract

Abstract: With the development of microblog, it is more convenient to comment on the Web. Up to now, there are very few studies on the sentiment classification for Chinese microblog, therefore this paper uses three machine learning algorithms, three kinds of feature selection methods and three feature weight methods to study the sentiment classification for Chinese microblog. The experimental results indicate that the performance of SVM is best in three machine learning algorithms, IG is the better feature selection method compared to the other methods, and TF-IDF is best fit for the sentiment classification in Chinese microblog. Combining the three factors the conclusion can be drawn that the performance of combination of SVM, IG and TF-IDF is best. For the movie domain it is found that the sentiment classification depends on the review style.

Key words: microblog, sentiment classification, machine learning, feature selection, term weight

摘要： 使用三种机器学习算法、三种特征选取算法以及三种特征项权重计算方法对微博进行了情感分类的实证研究。实验结果表明，针对不同的特征权重计算方法，支持向量机（SVM）和贝叶斯分类算法（NaIve Bayes）各有优势，信息增益（IG）特征选取方法相比于其他的方法效果明显要好。综合考虑三种因素，采用SVM和IG，以及TF-IDF（Term Frequency-Inverse Document Frequency）作为特征项权重，三者结合对微博的情感分类效果最好。针对电影领域，比较了微博评论和普通评论之间分类模型的通用性，实验结果表明情感分类性能依赖于评论的风格。

关键词: 微博, 情感分类, 机器学习, 特征选取, 特征项权重

LIU Zhiming, LIU Lu

. Empirical study of sentiment classification for Chinese microblog based on machine learning[J]. Computer Engineering and Applications, 2012, 48(1): 1-4.

刘志明，刘鲁. 基于机器学习的中文微博情感分类实证研究[J]. 计算机工程与应用, 2012, 48(1): 1-4.

[1]	RAN Rong, XU Xinghua, QIU Shaohua, CUI Xiaopeng, OUYANG Bin. Review of Crack Detection Methods Based on Deep Convolutional Neural Networks [J]. Computer Engineering and Applications, 2021, 57(9): 23-35.
[2]	YANG Chunxia, LI Xinxu, WU Jiajun, LIU Tianyu. Hierarchical Network Sentiment Classification Based on Attention Interaction Mechanism [J]. Computer Engineering and Applications, 2021, 57(9): 134-139.
[3]	ZHAO Yuanli, LIANG Zhijian. Research on Stance Detection Based on Dual Attention Mechanism of Heteronuclear Convolution [J]. Computer Engineering and Applications, 2021, 57(8): 119-125.
[4]	WEI Jihong, ZHENG Rongfeng, LIU Jiayong. Research on Malicious TLS Traffic Identification Based on Hybrid Neural Network [J]. Computer Engineering and Applications, 2021, 57(7): 107-114.
[5]	LI Li, JI Xinyuan, SONG Song. Prediction Model for Number of Software Defects in Loop [J]. Computer Engineering and Applications, 2021, 57(7): 158-163.
[6]	ZHANG Xiaoli, ZHANG Kuixing, JIANG Mei, WEI Benzheng, CONG Jinyu. Review of Image Classification Technology for Lymphoma [J]. Computer Engineering and Applications, 2021, 57(6): 1-9.
[7]	HAN Dongfang, Turdy Toheti, Askar Hamdulla. Survey on Question Classification Method in Question Answering System [J]. Computer Engineering and Applications, 2021, 57(6): 10-21.
[8]	LI Jingxing, YANG Youlong. Feature Selection of Markov Blanket for High Dimensional Data [J]. Computer Engineering and Applications, 2021, 57(6): 58-66.
[9]	WAN Mengxiang, YAO Hanbing. GAN Model for Malicious Web Training Data Generation [J]. Computer Engineering and Applications, 2021, 57(6): 124-130.
[10]	YANG Yemin, ZHANG Huijun, ZHANG Xiaolong. Research on Interpretable Visual Analysis Method of Random Forest [J]. Computer Engineering and Applications, 2021, 57(6): 168-175.
[11]	XU Kewen, XU Bo, WU Ying, XU Haoran. Overview of Application of Machine Learning in Ultrasound Images [J]. Computer Engineering and Applications, 2021, 57(4): 11-17.
[12]	WANG Zhendong, ZHANG Lin, LI Dahai. Survey of Intrusion Detection Systems for Internet of Things Based on Machine Learning [J]. Computer Engineering and Applications, 2021, 57(4): 18-27.
[13]	WANG Fang, ZHANG Xueying, HU Fengyun, LI Fenglian. Ensemble Method Classifies EEG from Stroke Patients [J]. Computer Engineering and Applications, 2021, 57(24): 276-282.
[14]	LYU Pin, WU Qinjuan, XU Jia. Intelligent Analysis of Text Information Disclosure of Listed Companies [J]. Computer Engineering and Applications, 2021, 57(24): 1-13.
[15]	ZHANG Yuxi, DUAN Zongtao, ZHU Yishui, WANG Luyang, ZHOU Yi, GUO Yu. Survey of Fuel Consumption Model for Motor Vehicle [J]. Computer Engineering and Applications, 2021, 57(24): 14-26.

Empirical study of sentiment classification for Chinese microblog based on machine learning

基于机器学习的中文微博情感分类实证研究

PDF

Knowledge

Abstract

Cite this article

share this article

References

Related Articles 15

Recommended Articles

Metrics