Abstract:Modern text-to-speech (TTS) systems are able to generate audio that sounds almost as natural as human speech. However, the bar of developing high-quality TTS systems remains high since a sizable set of studio-quality pairs is usually required. Compared to commercial data used to develop state-of-the-art systems, publicly available data are usually worse in terms of both quality and size. Audio generated by TTS systems trained on publicly available data tends to not only sound less natural, but also exhibits more background noise. In this work, we aim to lower TTS systems' reliance on high-quality data by providing them the textual knowledge extracted by deep pre-trained language models during training. In particular, we investigate the use of BERT to assist the training of Tacotron-2, a state of the art TTS consisting of an encoder and an attention-based decoder. BERT representations learned from large amounts of unlabeled text data are shown to contain very rich semantic and syntactic information about the input text, and have potential to be leveraged by a TTS system to compensate the lack of high-quality data. We incorporate BERT as a parallel branch to the Tacotron-2 encoder with its own attention head. For an input text, it is simultaneously passed into BERT and the Tacotron-2 encoder. The representations extracted by the two branches are concatenated and then fed to the decoder. As a preliminary study, although we have not found incorporating BERT into Tacotron-2 generates more natural or cleaner speech at a human-perceivable level, we observe improvements in other aspects such as the model is being significantly better at knowing when to stop decoding such that there is much less babbling at the end of the synthesized audio and faster convergence during training.

Knowledge Distillation from Bert in Pre-Training and Fine-Tuning for Polyphone Disambiguation.

A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese

Mandarin Text-to-Speech Front-End with Lightweight Distilled Convolution Network

g2pW: A Conditional Weighted Softmax BERT for Polyphone Disambiguation in Mandarin

Adaptive Knowledge Distillation between Text and Speech Pre-trained Models

Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning

External Knowledge Augmented Polyphone Disambiguation Using Large Language Model

MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models

Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers

Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-Speech

LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding

Knowledge Transfer from Pre-trained Language Models to Cif-based Speech Recognizers via Hierarchical Distillation

Towards Non-task-specific Distillation of BERT via Sentence Representation Approximation

Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models.

DDK: Dynamic structure pruning based on differentiable search and recursive knowledge distillation for BERT

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

Decouple Non-parametric Knowledge Distillation For End-to-end Speech Translation

One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers

Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling

Patient Knowledge Distillation for BERT Model Compression