| 講演抄録/キーワード |
| 講演名 |
2006-05-18 16:15
語義の違いを検出するための大規模コーパス処理手法の検討 ○相澤彰子(NII/総研大) |
| 抄録 |
(和) |
本稿では、タグなし自然言語文による大規模コーパスからの類語辞書自動構築法について検討する。まず、係り受け解析から得られる語の共起情報に基づき類語や例文を抽出するための手法の概要について述べる。次に、新聞記事コーパスを例にとり、コーパスが大規模になった場合の影響や同時クラスタリング法の効果を調べる。最後に実際にコーパスから構築した辞書の例を示す。 |
| (英) |
This paper focuses on issues in automatic extraction of synonyms from large scale untagged corpora. In the paper, a coocurrence analysis-based method is first introduced where synonyms and sample phrases are extracted simultaneously utilizing the result of word dependency analysis. Next, the influence of the corpus scale to the extraction result is examined using newspaper collections. A demonstrative example of the extracted dictionary is also shown. |
| キーワード |
(和) |
テキストコーパス / 類語辞書自動構築 / 語の共起情報 / テキストマイニング / / / / |
| (英) |
text corpora / automatic construction of synonymous words dictionaries / cooccurrencies of words / text mining / / / / |
| 文献情報 |
信学技報, vol. 106, no. 38, AI2006-11, pp. 57-62, 2006年5月. |
| 資料番号 |
AI2006-11 |
| 発行日 |
2006-05-11 (AI) |
| ISSN |
Print edition: ISSN 0913-5685 |
| PDFダウンロード |
|