| 講演抄録/キーワード |
| 講演名 |
2024-05-22 14:15
音声強調で音声認識性能はなぜ劣化するのか? ~ 音声強調誤差が音声認識性能に与える影響の分析 ~ ○落合 翼(NTT)・岩本一真(同志社大)・デルクロア マーク・池下林太郎・佐藤 宏・荒木章子(NTT)・片桐 滋(同志社大) EA2024-4 |
| 抄録 |
(和) |
深層学習技術は,シングルチャネル音声強調の音声強調性能を劇的に向上させた.しかし近年の研究において,こうしたシングルチャネル音声強調は音声認識性能の向上には必ずしも寄与せず,観測された雑音付き音声信号と比較してもむしろ音声認識性能を劣化させる場合もあることが多数報告されている.従来研究においては,シングルチャネル音声強調によって生じる処理歪みが音声認識を劣化させる要因であると仮定されていた.しかし,そうした音声認識性能を劣化させる処理歪みに関して,詳細な分析や数学的な指標は検討されていなかった.本研究では,シングルチャネル音声強調がなぜ音声認識性能を劣化させるのかについての分析を行い、音声強調誤差に含まれるartifact誤差 [1] が音声認識性能の劣化を引き起こす要因であることを明らかにした [2], [3].また,音声強調フロントエンドを音声認識レベルの学習基準で最適化することによりartifact誤差が低減されることを実験的に示すとともに,強調信号に観測信号を加えることでartifact誤差が低減できることを数理的,実験的に示した.こうした音声認識性能の劣化要因に関する知見は,音声認識に適したシングルチャネル音声強調フロントエンドの設計を考える上で重要なものである. |
| (英) |
Deep learning techniques have dramatically improved the speech enhancement (SE) performance of single-channel SE. However, recent studies has reported that such single-channel SE does not necessarily contribute to improve automatic speech recognition (ASR) performance but rather degrade it compared to inputting observed noisy signals directly. In conventional studies, it is often assumed that processing distortions induced by the single-channel SE are the cause that degrades ASR performance. However, there have been no detailed analyses and mathematical metrics to explain such processing distortions that could degrades the ASR performance. In this study, we analyzed why the single-channel SE degrades ASR performance, and revealed that artifact errors [1] contained in SE errors is the cause that degrades the ASR performance [2], [3]. We also showed experimentally that the artifact errors are reduced by optimizing the SE front-end based on the ASR-level training objective, and also showed mathematically and experimentally that the artifact errors can be reduced by adding an observation signal to the enhanced signal. Insights on this issue would be valuable for designing single-channel SE front-ends suitable for ASR. |
| キーワード |
(和) |
シングルチャネル音声強調 / ロバスト音声認識 / 処理歪み / / / / / |
| (英) |
single-channel speech enhancement / robust automatic speech recognition / processing distortion / / / / / |
| 文献情報 |
信学技報, vol. 124, no. 42, EA2024-4, pp. 20-21, 2024年5月. |
| 資料番号 |
EA2024-4 |
| 発行日 |
2024-05-15 (EA) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
EA2024-4 |