| 講演抄録/キーワード |
| 講演名 |
2009-01-26 15:00
依存構造から補強文脈自由文法の変換 ○田中穗積・中西良介(中京大) NLC2008-73 |
| 抄録 |
(和) |
自然言語処理技術には大別して、ルールベースの方法とコーパスベースの方法がある。わが国では、コーパスベースの方法が自然言語処理技術の研究の世界を席巻している。ルールベースの方法の最大の問題は、文法開発にある。本論文では、日本語の文に係り受け構造(依存構造)を付与した大量の既存のコーパスから、大規模な日本語の関数で補強された補強CFG(文脈自由文法形式の日本語文法)を機械的に抽出するアルゴリズムを提案する。この補強CFGを用いて日本語文を(入力文として文節列を与えて)既存のパーザでパーズすると構文木として関数木を得る。モンタギュー文法では、この関数木を分析木と呼んでいる。分析木は関数木の根の部分の関数を評価して、関数木沿った意味解釈を始める。こうして意味論と文法論とを一体化した自然言語処理を行うことができる。本論文では多種多様な文に対する依存構造付きのコーパスを大量に用意しておき、そこから大規模な日本語解析用文法(補強CFG)を機械的(自動的)に抽出するアルゴリズムを提案する。それにより、文法開発という、ルールベースの自然言語処理技術における最大の問題を解決する。 |
| (英) |
There are two technologies of the natural language processing. One is called a rule-based method and the other one, a corpus-based method. Although the rule-based method has many advantages, the development of large scale rule set with the wide coverage is very difficult and needs time consuming efforts. This is the reason why corpus-based method exceeds the rule-based method and becomes so popular in natural language processing. The author presents a new algorithm to extracts a grammar rule from the corpus with dependency structure. The algorithm constructs augmented CFG rules automatically and can generate large-scale rule set with wide coverage. The paper presents the details of the algorithm and points out the problems which have to solve in the future. |
| キーワード |
(和) |
コーパスベース / ルールベース / CFG / 関数木 / / / / |
| (英) |
Corpus-based / Rule-based / CFG / Function Tree / / / / |
| 文献情報 |
信学技報, vol. 108, no. 408, NLC2008-73, pp. 13-18, 2009年1月. |
| 資料番号 |
NLC2008-73 |
| 発行日 |
2009-01-19 (NLC) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
NLC2008-73 |