Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

The Iflytek System for Blizzard Machine Learning Challenge 2017-ES1

The USTC and iFlytek Speech Synthesis Systems for Blizzard Challenge 2007

The NLPR Speech Synthesis Entry for Blizzard Challenge 2020

DelightfulTTS: the Microsoft Speech Synthesis System for Blizzard Challenge 2021

BLSTM Guided Unit Selection Synthesis System for Blizzard Challenge 2016

The DeepZen Speech Synthesis System for Blizzard Challenge 2023

The Sogou Speech Synthesis System for Blizzard Challenge 2018

The USTC System for Blizzard Challenge 2009

Overview of NIT HMM-based speech synthesis system for Blizzard Challenge 2009

MuLanTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2023

The NTU-AISG Text-to-speech System for Blizzard Challenge 2020

The USTC System for Blizzard Challenge 2010

The USTC System for Blizzard Challenge 2008

USTC System for Blizzard Challenge 2006 an Improved HMM-based Speech Synthesis Method

The FruitShell French synthesis system at the Blizzard 2023 Challenge

The USTC System for Blizzard Machine Learning Challenge 2017-ES2

Analysis Syntactic Parsing Speech Synthesis Text Text Analysis Syntactic Parsing Speech Database Training Synthesis Training of HMMs Speech Database Data Selection Feature Extraction HMMs

The IMS Toucan System for the Blizzard Challenge 2023

The huya multi-speaker and multi-style speech synthesis system for m2voc challenge 2020

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.