| 講演抄録/キーワード |
| 講演名 |
2009-07-17 13:25
話者クラス音響モデルを用いた講演音声認識の性能向上 ○伊藤 貴・奥山洋平・加藤正治・小坂哲夫(山形大)・好田正紀(山形大名誉教授) SP2009-42 |
| 抄録 |
(和) |
本稿では講演音声認識の性能向上を目指し,話者クラス音響モデルの検討を行った.話者クラスモデルの使用法として,1)尤度基準によるモデルの自動選択,2)システム統合,の検討を行った.さらに,この認識結果を利用して教師なし適応の性能向上の検討を行った.以上の評価を日本語話し言葉コーパスを用いて行った.認識実験の結果,ベースラインの単語誤り率19.75\%に対し,話者クラスモデルの自動選択で19.11\%,システム統合で18.65\%を得た.また,一般的なMLLR 適応で17.50\%,話者クラス音響モデルを利用した適応で17.03\%,適応後の話者クラス音響モデルの出力統合により16.79\%を得た.以上より,講演音声認識において,提案手法が有効であることが分かった. |
| (英) |
This paper describes a new method based on speaker-class (SC) models in order to improve the performance of lecture speech recognition. We investigate two usages of SC models: 1) the automatic selection of SC
model by likelihood basis, and 2) the system combination of SC models. Furthermore, unsupervised speaker adaptation is studied by using SC models. The evaluation was conducted on CSJ (Corpus of Spontaneous Japanese). As the results, a word error rate of 19.11\% was obtained by using the automatic selection method, and 18.65\% was obtained by using the system combination, while 19.75\% was obtained in the baseline experiment. In addition, 17.03\% was obtained by using the adaptation method based on SC models, and 16.79\% was obtained by using the system combination based on adapted SC models, while 17.50\% was obtained by using conventional MLLR. The results showed that the proposed methods were effective for lecture speech recognition. |
| キーワード |
(和) |
大語彙連続音声認識 / 教師なし話者適応 / 話者クラスモデル / システム統合 / 日本語話し言葉コーパス / HMM / / |
| (英) |
LVCSR / unsupervised speaker adaptation / speaker-class model / system combination / Corpus of Spontaneous Japanese / HMM / / |
| 文献情報 |
信学技報, vol. 109, no. 139, SP2009-42, pp. 7-12, 2009年7月. |
| 資料番号 |
SP2009-42 |
| 発行日 |
2009-07-10 (SP) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
SP2009-42 |