| 講演抄録/キーワード |
| 講演名 |
2010-05-28 10:55
Webテキストにおける未知語の頻度調査 ○服部 峻・亀田弘之(東京工科大) TL2010-2 |
| 抄録 |
(和) |
日々増大して行くWebという情報源から様々な知識を抽出するWebマイニングの研究が盛んに行われているが,Webテキストを形態素解析や意味解析など自然言語処理する際,システムが用いる辞書に品詞や読み,意味などが未登録である「未知語」の存在が問題になる.本稿では,Webテキストに存在する多種なメディア,多様な話題,及び,投稿日時の3軸に依って,どのように未知語が分布しているか頻度調査を行った結果,Webテキストを自然言語処理するシステムにおいて,どんな分野で特に未知語処理が有用(必要)かなどの知見が得られたので報告する. |
| (英) |
Mining the Web to extract various knowledge from the growing source has become one of the hottest research topics. However, while such a Natural Language Processing (NLP) as morphological analysis or semantic analysis for Web text, the existence of ``Unknown Words'' that are not registered in a NLP system's dictionary (lexical database) is a serious impediment. In this paper, we survey the prevalence of unknown words in various domains of Japanese Web Text, e.g., dependency on its type of Web media, topics and upload date. |
| キーワード |
(和) |
未知語 / Web文書 / 未登録語 / 新語 / Webマイニング / 未知語処理 / 自然言語処理 / |
| (英) |
Unknown Words / Web Documents / Unregistered Words / New Words / Web Mining / Unknown Word Processing / Natural Language Processing / NLP |
| 文献情報 |
信学技報, vol. 110, no. 63, TL2010-2, pp. 7-12, 2010年5月. |
| 資料番号 |
TL2010-2 |
| 発行日 |
2010-05-21 (TL) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
TL2010-2 |