| 講演抄録/キーワード |
| 講演名 |
2011-07-07 16:30
日本語未知語のテキストからの自動獲得 ○村脇有吾・黒橋禎夫(京大) NLC2011-8 |
| 抄録 |
(和) |
日本語の形態素解析は,テキスト中に出現する形態素があらかじめ辞書に登録さ
れていることを前提としており,辞書に登録されていない未知語は解析誤りの
原因となっていた.
そのため,新たな分野のテキストを解析する際に,あらかじめ人手で形態素を追
加する必要があった.
この未知語問題を解決するために,我々はテキストから未知語を自動獲得し,人
手の 介在なしに語彙を増やして形態素解析を行うという研究を行なっている.
本稿では未知語の自動獲得の現状と課題を報告する. |
| (英) |
In Japanese morphological analysis, it is usually assumed that words in
text are listed in a pre-defined dictionary.
Errors are often caused by unknown words, or words not found in the
dictionary.
As a result, we need to register new words to the dictionary in advance
every time we are to process texts from a new domain.
To address this problem, we are working on a framework where unknown
words are automatically acquired from text and added to the dictionary
without manual supervision.
In this paper, we report recent progress and remaining problems in
unknown word acquisition. |
| キーワード |
(和) |
形態素解析 / 未知語 / 語彙獲得 / / / / / |
| (英) |
Japanese morphological analysis / unknown word / lexical acquisition / / / / / |
| 文献情報 |
信学技報, vol. 111, no. 119, NLC2011-8, pp. 37-42, 2011年7月. |
| 資料番号 |
NLC2011-8 |
| 発行日 |
2011-06-30 (NLC) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
NLC2011-8 |