Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

Text-Guided HuBERT: Self-Supervised Speech Pre-Training via Generative Adversarial Networks

Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.

Improved Speech Pre-Training with Supervision-Enhanced Acoustic Unit

Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture Speech

MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations

SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training

Unified Speech-Text Pre-training for Speech Translation and Recognition

PhonHuBERT: A Phoneme Transcription Tool for Song Datasets

SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

token2vec: A Joint Self-Supervised Pre-training Framework Using Unpaired Speech and Text

TESSP: Text-Enhanced Self-Supervised Speech Pre-training

Voice Deepfake Detection Using the Self-Supervised Pre-Training Model HuBERT

Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model

Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy

Fast-HuBERT: An Efficient Training Framework for Self-Supervised Speech Representation Learning

W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

HuBERTopic: Enhancing Semantic Representation of HuBERT through Self-supervision Utilizing Topic Model