高级检索

基于潜在语义索引的Web信息预测采集过滤方法

Forecast and Filter Method for Web Page Gathering Based on Latent Sematic Indexing

  • 摘要: Web信息急速膨胀使有效定向采集特定领域信息成为网上信息检索中一个日益重要的研究方向.提出一种基于潜在语义索引的Web信息预测采集过滤方法.在样本文档集潜在语义索引对文档相似计算的基础上,构造出用户兴趣模型,判断页面相关性进行文本过滤.通过对Web站点结构分析、对未知网页的相关性预测来控制信息采集过程.在保持定向采集精度的同时,缩短采集时间、减少存储、加快检索,节约了网络资源.

     

    Abstract: Following rapid expansion of huge information on Web, the efficient Web information gathering on specified fields becomes a key topic in information retrieval research. Based on LSI (Latent Semantic Indexing), this paper presents the forecast and filter method for Web page gathering. The method applies LSI on Web page sets provided by user to design the interested model and compute the relevance between Web page and the interested model for text filtering. Based on the analysis of Website structure, forecast for the relevance of Web page controls the gathering process. As a result, gathering time is shortened, storage decreased, retrieval speeded, and net resources are saved.

     

/

返回文章
返回