| Paper Abstract and Keywords |
| Presentation |
2023-02-28 09:30
MS-FC-HiFiGAN : Fast Neural Waveform Generation Model With Learnable Lightweight Upsampling Haruki Yamashita (Kobe Univ/NICT), Takuma Okamoto (NICT), Ryoichi Takashima, Tetsuya Takiguchi (Kobe Univ), Tomoki Toda (Nagoya Univ/NICT), Hisashi Kawai (NICT) EA2022-76 SIP2022-120 SP2022-40 |
| Abstract |
(in Japanese) |
(See Japanese page) |
| (in English) |
In recent years, in text-to-speech synthesis, it is required to improve the inference speed while keeping the quality.
Multi-stream(MS) iSTFT-HiFiGAN was proposed as a high-speed model of HiFi-GAN, a vocoder capable of inferring waveforms on single CPU.
In the TTS task using VITS, although there was some deterioration in sound quality, the speed was increased by about 4 times.
In this paper, we propose a MS-FC-HiFi-GAN in which the inverse short-time Fourier transform (iSTFT) part is changed to trainable fully connected layer for the purpose of improving the synthesis quality of the MS-iSTFT-HiFiGAN.
As for the inference speed, RTF was 0.15 on 1 CPU, which is the same as MS-iSTFT-HiFiGAN.
Synthesis quality was inferior to that of MS-iSTFT-HiFiGAN in TTS task, but was superior to thatin analysis/synthesis task. |
| Keyword |
(in Japanese) |
(See Japanese page) |
| (in English) |
speech synthesis / Neural Vocoder / HiFi-GAN / Text-to-Speech / Analysis Synthesis / / / |
| Reference Info. |
IEICE Tech. Rep., vol. 122, no. 389, SP2022-40, pp. 7-12, Feb. 2023. |
| Paper # |
SP2022-40 |
| Date of Issue |
2023-02-21 (EA, SIP, SP) |
| ISSN |
Online edition: ISSN 2432-6380 |
Copyright and reproduction |
All rights are reserved and no part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Notwithstanding, instructors are permitted to photocopy isolated articles for noncommercial classroom use without fee. (License No.: 10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| Download PDF |
EA2022-76 SIP2022-120 SP2022-40 |
| Conference Information |
| Committee |
SP IPSJ-SLP EA SIP |
| Conference Date |
2023-02-28 - 2023-03-01 |
| Place (in Japanese) |
(See Japanese page) |
| Place (in English) |
|
| Topics (in Japanese) |
(See Japanese page) |
| Topics (in English) |
|
| Paper Information |
| Registration To |
SP |
| Conference Code |
2023-02-SP-SLP-EA-SIP |
| Language |
Japanese |
| Title (in Japanese) |
(See Japanese page) |
| Sub Title (in Japanese) |
(See Japanese page) |
| Title (in English) |
MS-FC-HiFiGAN : Fast Neural Waveform Generation Model With Learnable Lightweight Upsampling |
| Sub Title (in English) |
|
| Keyword(1) |
speech synthesis |
| Keyword(2) |
Neural Vocoder |
| Keyword(3) |
HiFi-GAN |
| Keyword(4) |
Text-to-Speech |
| Keyword(5) |
Analysis Synthesis |
| Keyword(6) |
|
| Keyword(7) |
|
| Keyword(8) |
|
| 1st Author's Name |
Haruki Yamashita |
| 1st Author's Affiliation |
Kobe University/National Institute of Information and Communications Technology (Kobe Univ/NICT) |
| 2nd Author's Name |
Takuma Okamoto |
| 2nd Author's Affiliation |
National Institute of Information and Communications Technology (NICT) |
| 3rd Author's Name |
Ryoichi Takashima |
| 3rd Author's Affiliation |
Kobe University (Kobe Univ) |
| 4th Author's Name |
Tetsuya Takiguchi |
| 4th Author's Affiliation |
Kobe University (Kobe Univ) |
| 5th Author's Name |
Tomoki Toda |
| 5th Author's Affiliation |
Nagoya University/National Institute of Information and Communications Technology (Nagoya Univ/NICT) |
| 6th Author's Name |
Hisashi Kawai |
| 6th Author's Affiliation |
National Institute of Information and Communications Technology (NICT) |
| 7th Author's Name |
|
| 7th Author's Affiliation |
() |
| 8th Author's Name |
|
| 8th Author's Affiliation |
() |
| 9th Author's Name |
|
| 9th Author's Affiliation |
() |
| 10th Author's Name |
|
| 10th Author's Affiliation |
() |
| 11th Author's Name |
|
| 11th Author's Affiliation |
() |
| 12th Author's Name |
|
| 12th Author's Affiliation |
() |
| 13th Author's Name |
|
| 13th Author's Affiliation |
() |
| 14th Author's Name |
|
| 14th Author's Affiliation |
() |
| 15th Author's Name |
|
| 15th Author's Affiliation |
() |
| 16th Author's Name |
|
| 16th Author's Affiliation |
() |
| 17th Author's Name |
|
| 17th Author's Affiliation |
() |
| 18th Author's Name |
|
| 18th Author's Affiliation |
() |
| 19th Author's Name |
|
| 19th Author's Affiliation |
() |
| 20th Author's Name |
|
| 20th Author's Affiliation |
() |
| 21st Author's Name |
|
| 21st Author's Affiliation |
() |
| 22nd Author's Name |
|
| 22nd Author's Affiliation |
() |
| 23rd Author's Name |
|
| 23rd Author's Affiliation |
() |
| 24th Author's Name |
|
| 24th Author's Affiliation |
() |
| 25th Author's Name |
|
| 25th Author's Affiliation |
() |
| 26th Author's Name |
/ / |
| 26th Author's Affiliation |
()
() |
| 27th Author's Name |
/ / |
| 27th Author's Affiliation |
()
() |
| 28th Author's Name |
/ / |
| 28th Author's Affiliation |
()
() |
| 29th Author's Name |
/ / |
| 29th Author's Affiliation |
()
() |
| 30th Author's Name |
/ / |
| 30th Author's Affiliation |
()
() |
| 31st Author's Name |
/ / |
| 31st Author's Affiliation |
()
() |
| 32nd Author's Name |
/ / |
| 32nd Author's Affiliation |
()
() |
| 33rd Author's Name |
/ / |
| 33rd Author's Affiliation |
()
() |
| 34th Author's Name |
/ / |
| 34th Author's Affiliation |
()
() |
| 35th Author's Name |
/ / |
| 35th Author's Affiliation |
()
() |
| 36th Author's Name |
/ / |
| 36th Author's Affiliation |
()
() |
| Speaker |
Author-1 |
| Date Time |
2023-02-28 09:30:00 |
| Presentation Time |
20 minutes |
| Registration for |
SP |
| Paper # |
EA2022-76, SIP2022-120, SP2022-40 |
| Volume (vol) |
vol.122 |
| Number (no) |
no.387(EA), no.388(SIP), no.389(SP) |
| Page |
pp.7-12 |
| #Pages |
6 |
| Date of Issue |
2023-02-21 (EA, SIP, SP) |