Document image retrieval method using combination of text and non-text features

doi:10.3778/j.issn.1002-8331.2010.12.002

Computer Engineering and Applications ›› 2010, Vol. 46 ›› Issue (12): 5-8.DOI: 10.3778/j.issn.1002-8331.2010.12.002

• 博士论坛 • Previous Articles Next Articles

Document image retrieval method using combination of text and non-text features

ZHANG Tian

School of Information Science and Engineering，Shandong University，Jinan 250100，China

Received:2010-01-27 Revised:2010-03-12 Online:2010-04-21 Published:2010-04-21
Contact: ZHANG Tian

综合文字和非文字区域特征的文档图像检索

张田

山东大学信息科学与工程学院，济南 250100

通讯作者: 张田

Abstract

Abstract: An improved self-adaptive method for text area extraction is proposed.With it，the document image is segmented into text area and non-text area firstly.And then，for text area，local features and global features are extracted.The local features include gaps between connected characters，height and width of connected characters，and the global features contain writing style and paragraph features.For non-text area，the key block feature is extracted.After that，the retrieval method combines all the features to improve the accuracy.Meantime，multi-dimensional retrieval structure is introduced to improve the speed.The experiments performed on a large-scale document image database （including 12，024 images） reveal that the method is more efficient than existing ones.

Key words: document image retrieval, text area extraction, paragraph feature, multi-dimensional retrieval structure

摘要： 提出一种改进的自适应文字区域提取算法，将文档图像分割成文字区域和非文字区域。对文字区域提取连通字符间空白、连通字符高度和宽度等局部特征，以及书写样式、段落特征等全局特征；对非文字区域，提取关键块特征。然后利用检索算法将文字区域特征和非文字区域特征结合起来，提高检索的准确性。同时，在检索算法中引入多维数据检索结构，有效地提高检索速度。通过对大规模文档数据库（包含12 024个文档）的检索，表明该算法具有较高的效率，优于现有的一般文档图像检索算法。

关键词: 文档图像检索, 文字区域提取, 段落特征, 多维数据检索结构

CLC Number:

TP391

ZHANG Tian. Document image retrieval method using combination of text and non-text features[J]. Computer Engineering and Applications, 2010, 46(12): 5-8.

张田. 综合文字和非文字区域特征的文档图像检索[J]. 计算机工程与应用, 2010, 46(12): 5-8.

[1]	CHEN Wang¹，LI Bo1，SHI Yanjun²，TENG Hongfei². Differential evolution algorithm with estimation of distribution for solving RCPSP problem [J]. Computer Engineering and Applications, 2011, 47(4): 1-4.
[2]	SHA Quanyou¹，SHI Jinfa¹，QIN Xiansheng². Research on dynamical decomposition and optimization configuration in aeronautic manufacturing field [J]. Computer Engineering and Applications, 2011, 47(4): 9-12.
[3]	DAI Qin，LIU Jianbo，LIU Shibin. Analysis of remote sensing information extraction using swarm intelligence method [J]. Computer Engineering and Applications, 2011, 47(4): 13-16.
[4]	LIU Guangshuai，LI Bailin，HE Chaoming. Patch-graph sparse optimization methods based on piecewise smooth surfaces reconstruction [J]. Computer Engineering and Applications, 2011, 47(4): 22-25.
[5]	LONG Yinfang，SHANG Junna. Frequency offset estimation for MC-CDMA systems [J]. Computer Engineering and Applications, 2011, 47(4): 102-104.
[6]	YU Jiangde¹，WANG Xijie¹，FAN Xiaozhong². Comparing of importance of above-context versus below-context for Chinese word segmentation [J]. Computer Engineering and Applications, 2011, 47(4): 117-120.
[7]	PEI Yingbo¹，LIU Xiaoxia². Study on improved CHI for feature selection in Chinese text categorization [J]. Computer Engineering and Applications, 2011, 47(4): 128-130.
[8]	ZHANG Yu，LUO Ke. OC-SVM-based classification for large-scale data sets [J]. Computer Engineering and Applications, 2011, 47(4): 131-133.
[9]	LIU Ronghui^1，2，ZHENG Jianguo¹. Clustering algorithm in Deep Web based on Chinese word segmentation [J]. Computer Engineering and Applications, 2011, 47(4): 138-140.
[10]	CAI Rangjia. Tibetan studies of corpus description method [J]. Computer Engineering and Applications, 2011, 47(4): 146-148.
[11]	LIU Xiuling，LIU Jing，WANG Hongrui，GUO Lei. Fast collision detection based on improved honeycomb-shape spatial decomposition [J]. Computer Engineering and Applications, 2011, 47(4): 149-153.
[12]	ZHANG Cong，GUI Zhiguo. Non-linear image sharpening approach based on noise estimation [J]. Computer Engineering and Applications, 2011, 47(4): 154-156.
[13]	FU Xiaojun¹，GUO Pengjiang¹，GUO Jing²，FENG Jun². 3D model classification based on statistical features and Markov models [J]. Computer Engineering and Applications, 2011, 47(4): 157-159.
[14]	CHEN Huijie，LAI Huicheng，JIA Zhiqiang. Double color image information hiding based on image mix and wavelet transform [J]. Computer Engineering and Applications, 2011, 47(4): 171-173.
[15]	YANG Xiaoqin，JI Xiaoyong. Fast motion estimation algorithm based on H.264 [J]. Computer Engineering and Applications, 2011, 47(4): 174-175.

Document image retrieval method using combination of text and non-text features

综合文字和非文字区域特征的文档图像检索

PDF

Knowledge

Abstract

Cite this article

share this article

References

Related Articles 15

Recommended Articles

Metrics