| Paper Abstract and Keywords |
| Presentation |
2023-02-28 09:10
Comparison of fundamental frequency controllable fast neural waveform generative models. Sota Shimizu (Kobe Univ./NICT), Takuma Okamoto (NICT), Ryoichi Takashima, Tetsuya Takiguchi (Kobe Univ.), Tomoki Toda (Nagoya Univ./NICT), Hisashi Kawai (NICT) EA2022-75 SIP2022-119 SP2022-39 |
| Abstract |
(in Japanese) |
(See Japanese page) |
| (in English) |
Neural vocoders, which reconstruct speech waveforms from acoustic features with deep neural networks, have significantly improved synthetic speech quality compared to conventional source-filter vocoders. Neural vocoders, like conventional source filter vocoders, are required to be able to flexibly control attributes such as fundamental frequency ($f_{mathrm{o}}$). Harmonic-Net+ and SiFi-GAN have been proposed as models that can synthesize a speech waveform in real time on CPU, while maintaining controllability of $f_{mathrm{o}}$ and high synthetic speech quality. In this study, we conduct experiments to evaluate Harmonic-Net+, MS-Harmonic-Net+ and SiFi-GAN, which are fast neural vocoders with controllability of $f_{mathrm{o}}$ for unseen speaker synthesis. |
| Keyword |
(in Japanese) |
(See Japanese page) |
| (in English) |
speech synthesis / Neural vocoder / fundamental frequency control / real-time inference / / / / |
| Reference Info. |
IEICE Tech. Rep., vol. 122, no. 389, SP2022-39, pp. 1-6, Feb. 2023. |
| Paper # |
SP2022-39 |
| Date of Issue |
2023-02-21 (EA, SIP, SP) |
| ISSN |
Online edition: ISSN 2432-6380 |
Copyright and reproduction |
All rights are reserved and no part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Notwithstanding, instructors are permitted to photocopy isolated articles for noncommercial classroom use without fee. (License No.: 10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| Download PDF |
EA2022-75 SIP2022-119 SP2022-39 |
| Conference Information |
| Committee |
SP IPSJ-SLP EA SIP |
| Conference Date |
2023-02-28 - 2023-03-01 |
| Place (in Japanese) |
(See Japanese page) |
| Place (in English) |
|
| Topics (in Japanese) |
(See Japanese page) |
| Topics (in English) |
|
| Paper Information |
| Registration To |
SP |
| Conference Code |
2023-02-SP-SLP-EA-SIP |
| Language |
Japanese |
| Title (in Japanese) |
(See Japanese page) |
| Sub Title (in Japanese) |
(See Japanese page) |
| Title (in English) |
Comparison of fundamental frequency controllable fast neural waveform generative models. |
| Sub Title (in English) |
|
| Keyword(1) |
speech synthesis |
| Keyword(2) |
Neural vocoder |
| Keyword(3) |
fundamental frequency control |
| Keyword(4) |
real-time inference |
| Keyword(5) |
|
| Keyword(6) |
|
| Keyword(7) |
|
| Keyword(8) |
|
| 1st Author's Name |
Sota Shimizu |
| 1st Author's Affiliation |
Kobe University/National Institute of Information and Communications Technology (Kobe Univ./NICT) |
| 2nd Author's Name |
Takuma Okamoto |
| 2nd Author's Affiliation |
National Institute of Information and Communications Technology (NICT) |
| 3rd Author's Name |
Ryoichi Takashima |
| 3rd Author's Affiliation |
Kobe University (Kobe Univ.) |
| 4th Author's Name |
Tetsuya Takiguchi |
| 4th Author's Affiliation |
Kobe University (Kobe Univ.) |
| 5th Author's Name |
Tomoki Toda |
| 5th Author's Affiliation |
Nagoya University/National Institute of Information and Communications Technology (Nagoya Univ./NICT) |
| 6th Author's Name |
Hisashi Kawai |
| 6th Author's Affiliation |
National Institute of Information and Communications Technology (NICT) |
| 7th Author's Name |
|
| 7th Author's Affiliation |
() |
| 8th Author's Name |
|
| 8th Author's Affiliation |
() |
| 9th Author's Name |
|
| 9th Author's Affiliation |
() |
| 10th Author's Name |
|
| 10th Author's Affiliation |
() |
| 11th Author's Name |
|
| 11th Author's Affiliation |
() |
| 12th Author's Name |
|
| 12th Author's Affiliation |
() |
| 13th Author's Name |
|
| 13th Author's Affiliation |
() |
| 14th Author's Name |
|
| 14th Author's Affiliation |
() |
| 15th Author's Name |
|
| 15th Author's Affiliation |
() |
| 16th Author's Name |
|
| 16th Author's Affiliation |
() |
| 17th Author's Name |
|
| 17th Author's Affiliation |
() |
| 18th Author's Name |
|
| 18th Author's Affiliation |
() |
| 19th Author's Name |
|
| 19th Author's Affiliation |
() |
| 20th Author's Name |
|
| 20th Author's Affiliation |
() |
| 21st Author's Name |
|
| 21st Author's Affiliation |
() |
| 22nd Author's Name |
|
| 22nd Author's Affiliation |
() |
| 23rd Author's Name |
|
| 23rd Author's Affiliation |
() |
| 24th Author's Name |
|
| 24th Author's Affiliation |
() |
| 25th Author's Name |
|
| 25th Author's Affiliation |
() |
| 26th Author's Name |
/ / |
| 26th Author's Affiliation |
()
() |
| 27th Author's Name |
/ / |
| 27th Author's Affiliation |
()
() |
| 28th Author's Name |
/ / |
| 28th Author's Affiliation |
()
() |
| 29th Author's Name |
/ / |
| 29th Author's Affiliation |
()
() |
| 30th Author's Name |
/ / |
| 30th Author's Affiliation |
()
() |
| 31st Author's Name |
/ / |
| 31st Author's Affiliation |
()
() |
| 32nd Author's Name |
/ / |
| 32nd Author's Affiliation |
()
() |
| 33rd Author's Name |
/ / |
| 33rd Author's Affiliation |
()
() |
| 34th Author's Name |
/ / |
| 34th Author's Affiliation |
()
() |
| 35th Author's Name |
/ / |
| 35th Author's Affiliation |
()
() |
| 36th Author's Name |
/ / |
| 36th Author's Affiliation |
()
() |
| Speaker |
Author-1 |
| Date Time |
2023-02-28 09:10:00 |
| Presentation Time |
20 minutes |
| Registration for |
SP |
| Paper # |
EA2022-75, SIP2022-119, SP2022-39 |
| Volume (vol) |
vol.122 |
| Number (no) |
no.387(EA), no.388(SIP), no.389(SP) |
| Page |
pp.1-6 |
| #Pages |
6 |
| Date of Issue |
2023-02-21 (EA, SIP, SP) |