| Paper Abstract and Keywords |
| Presentation |
2025-03-02 09:30
Uncertainty-Based Streaming ASR with Evidential Deep Learning Hiroaki Sato, Asahi Sakuma, Ryuga Sugano, Tadashi Kumano, Yoshihiko Kawai (NHK STRL), Ogawa Tetsuji (Waseda Univ.) EA2024-77 SIP2024-112 SP2024-18 |
| Abstract |
(in Japanese) |
(See Japanese page) |
| (in English) |
We propose an attention-based encoder-decoder (AED) speech recognition model for streaming recognition using the evidential deep learning (EDL) framework. While AED models achieve high accuracy in offline speech recognition, they struggle with streaming recognition, outputting incorrect tokens before sufficient input is received. EDL models uncertainty using the Dirichlet distribution, addressing the limitation of softmax-based methods that do not account for output uncertainty. We introduce EDL into AED models to control token output timing based on uncertainty. During training, we use loss functions to decrease model uncertainty as more speech input is received. During inference, token output decisions are made by thresholding uncertainty values. Experiments on CSJ and Librispeech show that our method achieves superior streaming performance compared to existing approaches without modifying the AED model structure. We also show that error rate remains stable with varying speech input lengths, and that error rate and latency can be controlled by adjusting the threshold. |
| Keyword |
(in Japanese) |
(See Japanese page) |
| (in English) |
streaming ASR / evidential deep learning / attention-based encoder-decoder / / / / / |
| Reference Info. |
IEICE Tech. Rep., vol. 124, no. 391, SP2024-18, pp. 1-6, March 2025. |
| Paper # |
SP2024-18 |
| Date of Issue |
2025-02-23 (EA, SIP, SP) |
| ISSN |
Online edition: ISSN 2432-6380 |
Copyright and reproduction |
All rights are reserved and no part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Notwithstanding, instructors are permitted to photocopy isolated articles for noncommercial classroom use without fee. (License No.: 10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| Download PDF |
EA2024-77 SIP2024-112 SP2024-18 |
| Conference Information |
| Committee |
EA SIP SP IPSJ-SLP |
| Conference Date |
2025-03-02 - 2025-03-04 |
| Place (in Japanese) |
(See Japanese page) |
| Place (in English) |
|
| Topics (in Japanese) |
(See Japanese page) |
| Topics (in English) |
|
| Paper Information |
| Registration To |
SP |
| Conference Code |
2025-03-EA-SIP-SP-SLP |
| Language |
Japanese |
| Title (in Japanese) |
(See Japanese page) |
| Sub Title (in Japanese) |
(See Japanese page) |
| Title (in English) |
Uncertainty-Based Streaming ASR with Evidential Deep Learning |
| Sub Title (in English) |
|
| Keyword(1) |
streaming ASR |
| Keyword(2) |
evidential deep learning |
| Keyword(3) |
attention-based encoder-decoder |
| Keyword(4) |
|
| Keyword(5) |
|
| Keyword(6) |
|
| Keyword(7) |
|
| Keyword(8) |
|
| 1st Author's Name |
Hiroaki Sato |
| 1st Author's Affiliation |
NHK Science & Technology Research Laboratories (NHK STRL) |
| 2nd Author's Name |
Asahi Sakuma |
| 2nd Author's Affiliation |
NHK Science & Technology Research Laboratories (NHK STRL) |
| 3rd Author's Name |
Ryuga Sugano |
| 3rd Author's Affiliation |
NHK Science & Technology Research Laboratories (NHK STRL) |
| 4th Author's Name |
Tadashi Kumano |
| 4th Author's Affiliation |
NHK Science & Technology Research Laboratories (NHK STRL) |
| 5th Author's Name |
Yoshihiko Kawai |
| 5th Author's Affiliation |
NHK Science & Technology Research Laboratories (NHK STRL) |
| 6th Author's Name |
Ogawa Tetsuji |
| 6th Author's Affiliation |
Waseda University (Waseda Univ.) |
| 7th Author's Name |
|
| 7th Author's Affiliation |
() |
| 8th Author's Name |
|
| 8th Author's Affiliation |
() |
| 9th Author's Name |
|
| 9th Author's Affiliation |
() |
| 10th Author's Name |
|
| 10th Author's Affiliation |
() |
| 11th Author's Name |
|
| 11th Author's Affiliation |
() |
| 12th Author's Name |
|
| 12th Author's Affiliation |
() |
| 13th Author's Name |
|
| 13th Author's Affiliation |
() |
| 14th Author's Name |
|
| 14th Author's Affiliation |
() |
| 15th Author's Name |
|
| 15th Author's Affiliation |
() |
| 16th Author's Name |
|
| 16th Author's Affiliation |
() |
| 17th Author's Name |
|
| 17th Author's Affiliation |
() |
| 18th Author's Name |
|
| 18th Author's Affiliation |
() |
| 19th Author's Name |
|
| 19th Author's Affiliation |
() |
| 20th Author's Name |
|
| 20th Author's Affiliation |
() |
| 21st Author's Name |
|
| 21st Author's Affiliation |
() |
| 22nd Author's Name |
|
| 22nd Author's Affiliation |
() |
| 23rd Author's Name |
|
| 23rd Author's Affiliation |
() |
| 24th Author's Name |
|
| 24th Author's Affiliation |
() |
| 25th Author's Name |
|
| 25th Author's Affiliation |
() |
| 26th Author's Name |
/ / |
| 26th Author's Affiliation |
()
() |
| 27th Author's Name |
/ / |
| 27th Author's Affiliation |
()
() |
| 28th Author's Name |
/ / |
| 28th Author's Affiliation |
()
() |
| 29th Author's Name |
/ / |
| 29th Author's Affiliation |
()
() |
| 30th Author's Name |
/ / |
| 30th Author's Affiliation |
()
() |
| 31st Author's Name |
/ / |
| 31st Author's Affiliation |
()
() |
| 32nd Author's Name |
/ / |
| 32nd Author's Affiliation |
()
() |
| 33rd Author's Name |
/ / |
| 33rd Author's Affiliation |
()
() |
| 34th Author's Name |
/ / |
| 34th Author's Affiliation |
()
() |
| 35th Author's Name |
/ / |
| 35th Author's Affiliation |
()
() |
| 36th Author's Name |
/ / |
| 36th Author's Affiliation |
()
() |
| Speaker |
Author-1 |
| Date Time |
2025-03-02 09:30:00 |
| Presentation Time |
20 minutes |
| Registration for |
SP |
| Paper # |
EA2024-77, SIP2024-112, SP2024-18 |
| Volume (vol) |
vol.124 |
| Number (no) |
no.389(EA), no.390(SIP), no.391(SP) |
| Page |
pp.1-6 |
| #Pages |
6 |
| Date of Issue |
2025-02-23 (EA, SIP, SP) |