Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

Learning Deep and Wide Contextual Representations Using BERT for Statistical Parametric Speech Synthesis

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.

Prosody Modelling with Pre-trained Cross-utterance Representations for Improved Speech Synthesis

Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data

Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis

Improving Deep Neural Network Based Speech Synthesis Through Contextual Feature Parametrization and Multi-Task Learning

Statistical Parametric Speech Synthesis Using Generalized Distillation Framework

Speech BERT Embedding For Improving Prosody in Neural TTS

Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation

An Effective Contextual Language Modeling Framework for Speech Summarization with Augmented Features

Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language Understanding

Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis

Deep Belief Network-Based Post-Filtering For Statistical Parametric Speech Synthesis

The DeepZen Speech Synthesis System for Blizzard Challenge 2023

Extracting Spectral Features Using Deep Autoencoders with Binary Distributed Hidden Units for Statistical Parametric Speech Synthesis.

Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class Information

Modeling Spectral Envelopes Using Restricted Boltzmann Machines and Deep Belief Networks for Statistical Parametric Speech Synthesis

Deep Pre-Training Transformers for Scientific Paper Representation

ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading

Enhancing Prosodic Features by Adopting Pre-trained Language Model in Bahasa Indonesia Speech Synthesis

Spoken Style Learning with Multi-modal Hierarchical Context Encoding for Conversational Text-to-Speech Synthesis.