Mining peculiar compositions of frequent substrings from sparse text data using background texts

研究成果: 著書/レポートタイプへの貢献会議での発言

5 引用 (Scopus)

抜粋

We consider mining unusual patterns from text T. Unlike existing methods which assume probabilistic models and use simple estimation methods, we employ a set B of background text in addition to T and compositions w = xy of x and y as patterns. A string w is peculiar if there exist x and y such that w = xy, each of x and y is more frequent in B than in T, and conversely w = xy is more frequent in T. The frequency of xy in T is very small since x and y are infrequent in T, but xy is relatively abundant in T compared to xy in B. Despite these complex conditions for peculiar compositions, we develop a fast algorithm to find peculiar compositions using the suffix tree. Experiments using DNA sequences show scalability of our algorithm due to our pruning techniques and the superiority of the concept of the peculiar composition.

元の言語英語
ホスト出版物のタイトルMachine Learning and Knowledge Discovery in Databases - European Conference, ECML PKDD 2009, Proceedings
ページ596-611
ページ数16
エディションPART 1
DOI
出版物ステータス出版済み - 11 9 2009
イベントEuropean Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2009 - Bled, スロベニア
継続期間: 9 7 20099 11 2009

出版物シリーズ

名前Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
番号PART 1
5781 LNAI
ISSN(印刷物)0302-9743
ISSN(電子版)1611-3349

その他

その他European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2009
スロベニア
Bled
期間9/7/099/11/09

    フィンガープリント

All Science Journal Classification (ASJC) codes

  • Theoretical Computer Science
  • Computer Science(all)

これを引用

Ikeda, D., & Suzuki, E. (2009). Mining peculiar compositions of frequent substrings from sparse text data using background texts. : Machine Learning and Knowledge Discovery in Databases - European Conference, ECML PKDD 2009, Proceedings (PART 1 版, pp. 596-611). (Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics); 巻数 5781 LNAI, 番号 PART 1). https://doi.org/10.1007/978-3-642-04180-8_56