Abstract:A method for statistical parametric speech synthesis incorporating generative adversarial networks (GANs) is proposed. Although powerful deep neural networks (DNNs) techniques can be applied to artificially synthesize speech waveform, the synthetic speech quality is low compared with that of natural speech. One of the issues causing the quality degradation is an over-smoothing effect often observed in the generated speech parameters. A GAN introduced in this paper consists of two neural networks: a discriminator to distinguish natural and generated samples, and a generator to deceive the discriminator. In the proposed framework incorporating the GANs, the discriminator is trained to distinguish natural and generated speech parameters, while the acoustic models are trained to minimize the weighted sum of the conventional minimum generation loss and an adversarial loss for deceiving the discriminator. Since the objective of the GANs is to minimize the divergence (i.e., distribution difference) between the natural and generated speech parameters, the proposed method effectively alleviates the over-smoothing effect on the generated speech parameters. We evaluated the effectiveness for text-to-speech and voice conversion, and found that the proposed method can generate more natural spectral parameters and $F_0$ than conventional minimum generation error training algorithm regardless its hyper-parameter settings. Furthermore, we investigated the effect of the divergence of various GANs, and found that a Wasserstein GAN minimizing the Earth-Mover's distance works the best in terms of improving synthetic speech quality.

D-MelGAN: speech synthesis with specific voiceprint features

SE-MelGAN -- Speaker Agnostic Rapid Speech Enhancement

Chinese personalised text‐to‐speech synthesis for robot human–machine interaction

Audio-driven Talking Face Video Generation with Natural Head Pose

MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Statistical parametric speech synthesis using generative adversarial networks under a multi-task learning framework

SingGAN: Generative Adversarial Network for High-Fidelity Singing Voice Generation

Multi-Band Melgan: Faster Waveform Generation For High-Quality Text-To-Speech

MixGAN-TTS: Efficient and Stable Speech Synthesis Based on Diffusion Model

Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network

Multi-band melgan: fasterwaveform generation for high-quality text-to-speech

On the Use of Audio Fingerprinting Features for Speech Enhancement with Generative Adversarial Network

MelGAN-VC: Voice Conversion and Audio Style Transfer on arbitrarily long samples using Spectrograms

DSPGAN: a GAN-based universal vocoder for high-fidelity TTS by time-frequency domain supervision from DSP

Using Speech Enhancement to Realize Speech Synthesis of Low-Resource Dungan Languages

Voice Conversion with Denoising Diffusion Probabilistic GAN Models

HiFi-WaveGAN: Generative Adversarial Network with Auxiliary Spectrogram-Phase Loss for High-Fidelity Singing Voice Generation

Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks

High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram