计算机工程与应用 ›› 2008, Vol. 44 ›› Issue (9): 143-146.

• 数据库、信号与信息处理 • 上一篇    下一篇

基于网页内容块策略的主题爬行

吴晓平,张长利,朱丽娜   

  1. 沈阳炮兵学院 基础部计算机实验中心,沈阳 110162

  • 收稿日期:2007-07-12 修回日期:2007-10-24 出版日期:2008-03-21 发布日期:2008-03-21
  • 通讯作者: 吴晓平

Block-based topic crawling

WU Xiao-ping,ZHANG Chang-li,ZHU Li-na

  

  1. Computer Experiment Center,Shenyang Artillery College,Shenyang 110162,China
  • Received:2007-07-12 Revised:2007-10-24 Online:2008-03-21 Published:2008-03-21
  • Contact: WU Xiao-ping

摘要: 因特网的迅速发展对传统的爬行器和搜索引擎提出了巨大的挑战。各种针对特定领域、特定人群的搜索引擎应运而生。Web主题信息搜索系统(网络蜘蛛)是主题搜索引擎的最主要的部分,它的任务是将搜集到的符合要求的Web页面返回给用户或保存在索引库中。Web 上的信息资源如此广泛,如何全面而高效地搜集到感兴趣的内容是网络蜘蛛的研究重点。提出了基于网页分块技术的主题爬行,实验结果表明,相对于其它的爬行算法,提出的算法具有较高的效率、爬准率、爬全率及穿越隧道的能力。

关键词: 定题搜索, 主题爬行, 搜索引擎, 爬行算法, 相关度分析

Abstract: With the explosive growth of the World-Wide Web,to general-purpose crawlers and search engines which pose great challenges.All sorts of special topic search engines are designed for special people and special domains.The web topic information search system(web spider) is the most important part of topic search engine,it collects web pages of special topic and provides users with the result or stores it in index database.Information resource of web is so extensive,how to collect interest content comprehensively and effectively,it is important to web spider research.In this paper,a new crawling strategy block-based topic crawling has been proposed,the experiments show that compared with some traditional algorithms,this algorithm has better performance.It is effective and has high precision.

Key words: topic-specific search, topic crawling, search engine, crawling algorithm, correlation analysis