| 講演抄録/キーワード |
| 講演名 |
2023-06-23 13:50
暗黙的言語情報を置換するCTCデコーダを用いた ストリーミング音声認識 ○高城巽成(豊橋技科大)・小川厚徳(NTT)・北岡教英・若林佑幸(豊橋技科大) SP2023-12 |
| 抄録 |
(和) |
音声認識技術は, 音声認識モデルの精度向上により, さまざまな分野で利用されているが, 学習に用いるデータと認識対象となるデータのドメインが異なる場合, 認識精度が低下する. この問題を解決する方法として, 大量のテキストで学習された言語モデルを用いる様々な手法が提案されている. 近年では, 音声認識モデルと言語モデルの統合方法として Shallow Fusion を拡張した Density Ratio Approach (DRA) が提案されている. しかしながら, 日本語音声において CTC デコーダを用いたストリーミング可能な音声認識モデルでの DRA の適用は, 未だに検討されていない. そこで本研究では, CTC デコーダを用いたストリーミング音声認識にて DRA によるドメイン適応を行った. ストリーミング処理を可能とするために, デコーダにてフレーム単位で逐次的に言語情報を置換していくことで, greedy search による認識結果を得る. また, CTC の条件付き独立性の仮定について考慮し, 置換する言語情報を選択した. 実験の結果から提案手法を用いることで認識精度が向上することを示した. |
| (英) |
Speech recognition technology has been employed in various fields due to the enhancement of speech recognition model accuracy. However, when the domain of the data used for training differs from that of the data to be recognized, recognition accuracy declines. To address this issue, several approaches utilizing language models trained on extensive text data have been proposed. Recently, the Density Ratio Approach (DRA), an extension of Shallow Fusion, has been introduced as a method for integrating speech recognition models with language models. Nevertheless, the application of DRA to streaming speech recognition models using a CTC decoder for Japanese speech has not been investigated. In this study, we conducted domain adaptation using DRA in streaming speech recognition with a CTC decoder. To facilitate streaming processing, the decoder successively replaces linguistic information on a frame-by-frame basis, obtaining recognition results through greedy search. Furthermore, we selected the linguistic information to be replaced, considering the assumption of conditional independence of
CTC. Experimental results demonstrate that the proposed method enhances recognition accuracy. |
| キーワード |
(和) |
End-to-End音声認識 / ストリーミング音声認識 / 言語モデル / CTC / DRA / / / |
| (英) |
End-to-End Speech Recognition / Streaming Speech Recognition / Language Mode / CTC / DRA / / / |
| 文献情報 |
信学技報, vol. 123, no. 88, SP2023-12, pp. 60-64, 2023年6月. |
| 資料番号 |
SP2023-12 |
| 発行日 |
2023-06-16 (SP) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
SP2023-12 |
| 研究会情報 |
| 研究会 |
SP IPSJ-MUS IPSJ-SLP |
| 開催期間 |
2023-06-23 - 2023-06-24 |
| 開催地(和) |
電気通信大学 |
| 開催地(英) |
|
| テーマ(和) |
音学シンポジウム2023 |
| テーマ(英) |
|
| 講演論文情報の詳細 |
| 申込み研究会 |
SP |
| 会議コード |
2023-06-SP-MUS-SLP |
| 本文の言語 |
日本語 |
| タイトル(和) |
暗黙的言語情報を置換するCTCデコーダを用いた ストリーミング音声認識 |
| サブタイトル(和) |
|
| タイトル(英) |
Streaming End-to-End speech recognition using a CTC decoder with substituted linguistic information |
| サブタイトル(英) |
|
| キーワード(1)(和/英) |
End-to-End音声認識 / End-to-End Speech Recognition |
| キーワード(2)(和/英) |
ストリーミング音声認識 / Streaming Speech Recognition |
| キーワード(3)(和/英) |
言語モデル / Language Mode |
| キーワード(4)(和/英) |
CTC / CTC |
| キーワード(5)(和/英) |
DRA / DRA |
| キーワード(6)(和/英) |
/ |
| キーワード(7)(和/英) |
/ |
| キーワード(8)(和/英) |
/ |
| 第1著者 氏名(和/英/ヨミ) |
高城 巽成 / Tatsunari Takagi / タカギ タツナリ |
| 第1著者 所属(和/英) |
豊橋技術科学大学 (略称: 豊橋技科大)
Toyohashi Univerdity of Technology (略称: TUT) |
| 第2著者 氏名(和/英/ヨミ) |
小川 厚徳 / Atsunori Ogawa / オガワ アツノリ |
| 第2著者 所属(和/英) |
日本電信電話株式会社 (略称: NTT)
NIPPON TELEGRAPH AND TELEPHONE CORPORATION (略称: NTT) |
| 第3著者 氏名(和/英/ヨミ) |
北岡 教英 / Norihide Kitaoka / キタオカ ノリヒデ |
| 第3著者 所属(和/英) |
豊橋技術科学大学 (略称: 豊橋技科大)
Toyohashi Univerdity of Technology (略称: TUT) |
| 第4著者 氏名(和/英/ヨミ) |
若林 佑幸 / Yukoh Wakabayashi / ワカバヤシ ユウコウ |
| 第4著者 所属(和/英) |
豊橋技術科学大学 (略称: 豊橋技科大)
Toyohashi Univerdity of Technology (略称: TUT) |
| 第5著者 氏名(和/英/ヨミ) |
/ / |
| 第5著者 所属(和/英) |
(略称: )
(略称: ) |
| 第6著者 氏名(和/英/ヨミ) |
/ / |
| 第6著者 所属(和/英) |
(略称: )
(略称: ) |
| 第7著者 氏名(和/英/ヨミ) |
/ / |
| 第7著者 所属(和/英) |
(略称: )
(略称: ) |
| 第8著者 氏名(和/英/ヨミ) |
/ / |
| 第8著者 所属(和/英) |
(略称: )
(略称: ) |
| 第9著者 氏名(和/英/ヨミ) |
/ / |
| 第9著者 所属(和/英) |
(略称: )
(略称: ) |
| 第10著者 氏名(和/英/ヨミ) |
/ / |
| 第10著者 所属(和/英) |
(略称: )
(略称: ) |
| 第11著者 氏名(和/英/ヨミ) |
/ / |
| 第11著者 所属(和/英) |
(略称: )
(略称: ) |
| 第12著者 氏名(和/英/ヨミ) |
/ / |
| 第12著者 所属(和/英) |
(略称: )
(略称: ) |
| 第13著者 氏名(和/英/ヨミ) |
/ / |
| 第13著者 所属(和/英) |
(略称: )
(略称: ) |
| 第14著者 氏名(和/英/ヨミ) |
/ / |
| 第14著者 所属(和/英) |
(略称: )
(略称: ) |
| 第15著者 氏名(和/英/ヨミ) |
/ / |
| 第15著者 所属(和/英) |
(略称: )
(略称: ) |
| 第16著者 氏名(和/英/ヨミ) |
/ / |
| 第16著者 所属(和/英) |
(略称: )
(略称: ) |
| 第17著者 氏名(和/英/ヨミ) |
/ / |
| 第17著者 所属(和/英) |
(略称: )
(略称: ) |
| 第18著者 氏名(和/英/ヨミ) |
/ / |
| 第18著者 所属(和/英) |
(略称: )
(略称: ) |
| 第19著者 氏名(和/英/ヨミ) |
/ / |
| 第19著者 所属(和/英) |
(略称: )
(略称: ) |
| 第20著者 氏名(和/英/ヨミ) |
/ / |
| 第20著者 所属(和/英) |
(略称: )
(略称: ) |
| 第21著者 氏名(和/英/ヨミ) |
/ / |
| 第21著者 所属(和/英) |
(略称: )
(略称: ) |
| 第22著者 氏名(和/英/ヨミ) |
/ / |
| 第22著者 所属(和/英) |
(略称: )
(略称: ) |
| 第23著者 氏名(和/英/ヨミ) |
/ / |
| 第23著者 所属(和/英) |
(略称: )
(略称: ) |
| 第24著者 氏名(和/英/ヨミ) |
/ / |
| 第24著者 所属(和/英) |
(略称: )
(略称: ) |
| 第25著者 氏名(和/英/ヨミ) |
/ / |
| 第25著者 所属(和/英) |
(略称: )
(略称: ) |
| 第26著者 氏名(和/英/ヨミ) |
/ / |
| 第26著者 所属(和/英) |
(略称: )
(略称: ) |
| 第27著者 氏名(和/英/ヨミ) |
/ / |
| 第27著者 所属(和/英) |
(略称: )
(略称: ) |
| 第28著者 氏名(和/英/ヨミ) |
/ / |
| 第28著者 所属(和/英) |
(略称: )
(略称: ) |
| 第29著者 氏名(和/英/ヨミ) |
/ / |
| 第29著者 所属(和/英) |
(略称: )
(略称: ) |
| 第30著者 氏名(和/英/ヨミ) |
/ / |
| 第30著者 所属(和/英) |
(略称: )
(略称: ) |
| 第31著者 氏名(和/英/ヨミ) |
/ / |
| 第31著者 所属(和/英) |
(略称: )
(略称: ) |
| 第32著者 氏名(和/英/ヨミ) |
/ / |
| 第32著者 所属(和/英) |
(略称: )
(略称: ) |
| 第33著者 氏名(和/英/ヨミ) |
/ / |
| 第33著者 所属(和/英) |
(略称: )
(略称: ) |
| 第34著者 氏名(和/英/ヨミ) |
/ / |
| 第34著者 所属(和/英) |
(略称: )
(略称: ) |
| 第35著者 氏名(和/英/ヨミ) |
/ / |
| 第35著者 所属(和/英) |
(略称: )
(略称: ) |
| 第36著者 氏名(和/英/ヨミ) |
/ / |
| 第36著者 所属(和/英) |
(略称: )
(略称: ) |
| 講演者 |
第1著者 |
| 発表日時 |
2023-06-23 13:50:00 |
| 発表時間 |
140分 |
| 申込先研究会 |
SP |
| 資料番号 |
SP2023-12 |
| 巻番号(vol) |
vol.123 |
| 号番号(no) |
no.88 |
| ページ範囲 |
pp.60-64 |
| ページ数 |
5 |
| 発行日 |
2023-06-16 (SP) |
|