| 講演抄録/キーワード |
| 講演名 |
2025-05-28 15:55
[招待講演]機械学習を用いた音響イベント検知・定位技術の実環境適用に向けた検討 ○安田昌弘・原田 登・齊藤翔一郎(NTT) EA2025-10 |
| 抄録 |
(和) |
音響イベント検知・定位 (SELD) は,多チャネル音響信号から周囲環境で発生した音響イベントの時刻・種類・方向を同時に推定する技術であり,歩行者安全支援やイマーシブ・コミュニケーション等への応用が期待されている.本稿では,SELDシステムの実環境適用に向け,大きく二つの課題に取り組むものである.
第一の課題は,センサ自体が自己運動する状況でのSELDシステムの音源定位性能が低下することである.この課題に取り組むため,我々は,頭部の6自由度動作と18ch音響信号を同時計測した新規データセットを収録した.加えて,音響信号と頭部運動情報の両方を利用したマルチモーダルSELDシステムを提案する.
第二の課題は,SELDシステムの学習に必要な音響イベントの位置についてのアノテーション収集のコストの高さに起因する,実環境学習データの不足である.この課題に対し,我々は音響イベントの位置についてのラベルを必要としない新しいSELDシステムの学習方法を提案する. |
| (英) |
Sound Event Localization and Detection (SELD) estimates the timing, class, and direction of sound events from multichannel audio, enabling applications such as pedestrian safety support and immersive communication. To advance SELD in real-world scenarios, we address two challenges. First, self-motion of the microphone array degrades localization performance; we record a new dataset with 18-channel audio and full six-degree-of-freedom head tracking and develop a multimodal model that fuses inertial and acoustic cues. Second, collecting spatial annotations is costly; we introduce a novel training framework that dispenses with spatial annotations by treating beamformed signals as instances in a weakly supervised learning scheme. |
| キーワード |
(和) |
音響イベント検知 / 音源定位 / 深層学習 / 6自由度 / 弱教師有り学習 / / / |
| (英) |
sound event detection / source localization / deep learning / six degree of freedom / weekly supervised learning / / / |
| 文献情報 |
信学技報, vol. 125, no. 36, EA2025-10, pp. 43-48, 2025年5月. |
| 資料番号 |
EA2025-10 |
| 発行日 |
2025-05-21 (EA) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
EA2025-10 |