| 講演抄録/キーワード |
| 講演名 |
2024-03-14 14:00
複数人対話環境における発話推定の学習データの組み合わせと分割数に関する考察 ○上村海斗・堀尾恵一(九工大) SIS2023-48 |
| 抄録 |
(和) |
今日,会議・ニュース・電話音声などを主な対象として話者ダイアライゼーションと呼ばれる発話区間検出技術の重要性が増してきている.しかし,従来のニューラルネットワークを用いた話者ダイアライゼーションには膨大な学習データを前提としている.本研究では同性の発話者2人の発話音声をそれぞれ録音し,それらを分割して組み合わせることで合成音声を作成することで学習データを生成した.分割数と組み合わせがテストデータに対する精度に与える影響を検証した. |
| (英) |
Today, the importance of a speech segment detection technique called speaker diarization is increasing, mainly in the fields of conference, news, and telephone speech. However, conventional neural network-based speaker diarization assumes a large amount of training data. In this study, training data was generated by recording the speech of two speakers of the same gender and creating a synthetic voice by splitting and combining them. The effect of the number of segmentations and combinations on the accuracy of the test data was verified. |
| キーワード |
(和) |
発話区間検出 / 双方向LSTM / 音声認識 / / / / / |
| (英) |
Diarization / Bidirectional LSTM / voice recognition / / / / / |
| 文献情報 |
信学技報, vol. 123, no. 440, SIS2023-48, pp. 17-20, 2024年3月. |
| 資料番号 |
SIS2023-48 |
| 発行日 |
2024-03-07 (SIS) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
SIS2023-48 |