Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

Research on Speech Synthesis Based on Mixture Alignment Mechanism

End-to-end Code-switched TTS with Mix of Monolingual Recordings.

MoBoAligner: A Neural Alignment Model for Non-Autoregressive TTS with Monotonic Boundary Search.

MixGAN-TTS: Efficient and Stable Speech Synthesis Based on Diffusion Model

HMM Based TTS for Mixed Language Text.

DIA-TTS: Deep-Inherited Attention-Based Text-to-Speech Synthesizer

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models

Initial investigation of an encoder-decoder end-to-end TTS framework using marginalization of monotonic hard latent alignments

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.

One TTS Alignment To Rule Them All

Neural Speech Synthesis with Transformer Network.

A Novel Hybrid Mandarin Speech Synthesis System Using Different Base Units for Model Training and Concatenation

An initial research: Towards accurate pitch extraction for speech synthesis based on BLSTM

Prosody Learning Mechanism for Speech Synthesis System Without Text Length Limit

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

Representation Mixing for TTS Synthesis

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Robust ASR

QS-TTS: Towards Semi-Supervised Text-to-Speech Synthesis via Vector-Quantized Self-Supervised Speech Representation Learning

AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling