基于本体论的Web信息抽取
Ontology-Based Information Extraction from Web Sources
-
摘要: 以本体论为基础,以所要提取的信息的层次结构作为信息提取的路径,定义了Web页面的信息项本体,并自动解析生成Web页面的结构本体.通过对这两个本体进行对比,构造了一种归纳学习算法来半自动地生成信息提取规则,对Web页面的信息提取具有较高的效率.Abstract: bstract Based on the ontologythis paper regards the hiberarchy of information to be extracted as the path of information extractiondefines an information item ontology of Web page and automatic creates a construction ontology by parsing the Web page.Using these two ontologiesa novel approach to sem-i automatically generate information extraction rules is presented for efficiently collecting information from Web.
下载: