| 講演抄録/キーワード |
| 講演名 |
2009-01-27 10:30
類似度の高いサブクラスタに基づく名詞クラスタリング ○金本聖也・竹内孔一(岡山大) NLC2008-76 |
| 抄録 |
(和) |
Pantel らがCBC という類似度の高いサブクラスタをあらかじめ作成しておく事でサブクラスタに基づいた揺れの少ない統合と語義を考慮した再統合を行うクラスタリング手法を提案したが,本研究ではCBC を基に係り受けパターンを利用した名詞クラスタリングを行い同義語・類義語クラスタの獲得を目指す.本論文ではCBC の既存の式ではなく確率分布を用いた類似度計算式(Jensen-Shannon) の使用,並びにサブクラスタ候補を決定する新しいスコアリング方法を用いた日本語の名詞クラスタリング手法を提案する.毎日新聞94 年度1 年分を用いてCBCに用いられる類似度計算式とJensen-Shannon の比較を行いJensen-Shannon の有効性を示し,さらにスコアリング式をいくつかのパターンで提案・比較を行い適切にサブクラスタ候補を決定するスコアリング方法を求める. |
| (英) |
In this paper we propose a noun clustering approach on the basis of CBC proposed by Pantel. CBC
is a clustering approach that carefully extracts clusters by finding sub-clusters regarded as committees with the
same meanings, and try to extract unknown clusters from the remaining elements. In preliminary experiments ofJapanese noun clustering, however, we found that CBC does not work well at (1) the measurement of basic similarity between words with context vectors and (2) scoring method that decides to merge sub-clusters. To these problems in this paper we propose to apply Jensen-Shannon formula as a measurement and a new scoring method. In the experimental results of constructing sub-clusters of Japanese nouns from a new paper article we will show that our proposed approaches overcome the approaches in CBC at the clustering accuracy. |
| キーワード |
(和) |
クラスタリング / 同義語 / 類義語 / / / / / |
| (英) |
Clustering / Synonym / Hypernymy / Hyponymy / / / / |
| 文献情報 |
信学技報, vol. 108, no. 408, NLC2008-76, pp. 31-35, 2009年1月. |
| 資料番号 |
NLC2008-76 |
| 発行日 |
2009-01-19 (NLC) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
NLC2008-76 |