| 講演抄録/キーワード |
| 講演名 |
2019-02-23 11:00
日本語文の係り受け木からの意味構造パターンマイニング ○鈴木 諒・兼岩 憲(電通大) AI2018-46 |
| 抄録 |
(和) |
Web上の大量テキストに対して,ユーザはキーワードによるWeb検索サービスから目的の情報を得ることができる.しかし,構造的なデータとして解読できないテキストからキーワード検索しても意味的にマッチした情報を探すには限界がある.本研究では,Web上の大規模なテキストに対して,日本語の自由な語順からでも意味構造を処理できる係り受け木の特徴表現(SITリスト)とその頻出パターンの抽出方法を提案する.評価実験では,SITリストを用いることで,日本語文の係り受け木に内在する共通の意味構造パターンを獲得できることを示す. |
| (英) |
Users can obtain information from large-scale web texts using keyword-based web search services. Such search services find information including keywords but not semantically matching information because natural language texts on the web are not machine-readable. In this paper, we develop a pattern mining method that extracts frequent semantic structures in the dependency trees of Japanese sentences. In order to deal with the flexible word order in Japanese, we propose feature expressions (SIT lists) consisting of the three sets of phase nodes in the dependency trees. In the evaluation experiment, we show that the pattern mining method for SIT lists enables us to flexibly extract common semantic patterns that are implicitly included in the dependency trees of Japanese sentences. |
| キーワード |
(和) |
テキストマイニング / 情報抽出 / 知識獲得 / / / / / |
| (英) |
Text mining / Information extraction / Knowledge acquisition / / / / / |
| 文献情報 |
信学技報, vol. 118, no. 453, AI2018-46, pp. 51-55, 2019年2月. |
| 資料番号 |
AI2018-46 |
| 発行日 |
2019-02-15 (AI) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
AI2018-46 |