Extraction of relevant snippets from web pages using hybrid features

Jun Zeng, Junhao Wen, Qingyu Xiong, Sachio Hirokawa

    研究成果: Chapter in Book/Report/Conference proceedingConference contribution

    抄録

    As the amount of web pages increase, identifying and retrieving distinct contents from the web has increasingly become more and more difficult. The traditional approach for extracting data from web page documents is to analyze the DOM (Document Object Model) structure of a HTML page and find a common pattern. However, the number of possible DOM layout patterns is virtually infinite, which means that there is no common pattern that can be used for all kinds of web pages. In this paper, we focus on the pages that are linked to a search engine and aim to analyze the features of relevant and meaningful contents instead of a common pattern. Three features of relevant snippets are introduced. They are: quantity of text, correlation between snippet and query that is inputted into a search engine, and HTML structure. Nine parameters are used to describe the three features. Also, a SVM learning experiment is conducted to verify the effectiveness of the three features. The results show that the HTML structure feature is the most effective feature which can determine whether a snippet is relevant or not.

    本文言語英語
    ホスト出版物のタイトルProceedings of the 2012 IIAI International Conference on Advanced Applied Informatics, IIAIAAI 2012
    ページ209-213
    ページ数5
    DOI
    出版ステータス出版済み - 12 14 2012
    イベント1st IIAI International Conference on Advanced Applied Informatics, IIAIAAI 2012 - Fukuoka, 日本
    継続期間: 9 20 20129 22 2012

    出版物シリーズ

    名前Proceedings of the 2012 IIAI International Conference on Advanced Applied Informatics, IIAIAAI 2012

    その他

    その他1st IIAI International Conference on Advanced Applied Informatics, IIAIAAI 2012
    Country日本
    CityFukuoka
    Period9/20/129/22/12

    All Science Journal Classification (ASJC) codes

    • Information Systems

    フィンガープリント 「Extraction of relevant snippets from web pages using hybrid features」の研究トピックを掘り下げます。これらがまとまってユニークなフィンガープリントを構成します。

    引用スタイル