Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition

An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition.

Improving Automatic Speech Recognition Performance for Low-Resource Languages With Self-Supervised Models

SUPERB: Speech Understanding and PERformance Benchmark

Progressive Residual Extraction based Pre-training for Speech Representation Learning

Improving Hybrid CTC/Attention End-to-end Speech Recognition with Pretrained Acoustic and Language Model

Utilizing Self-supervised Representations for MOS Prediction

Large-Scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

Semi-Supervised Spoken Language Understanding Via Self-Supervised Speech and Language Model Pretraining.

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.

Exploring Self-supervised Pre-trained ASR Models For Dysarthric and Elderly Speech Recognition

Large-Scale Unsupervised Pre-Training for End-to-End Spoken Language Understanding.

BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition

Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention Networks

SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning

Investigating Self-supervised Pretraining Frameworks for Pathological Speech Recognition

Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition

Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information

Progressive Multi-scale Self-supervised Learning for Speech Recognition