Abstract:Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data. However, the lack of aligned data poses a major practical problem for TTS and ASR on low-resource languages. In this paper, by leveraging the dual nature of the two tasks, we propose an almost unsupervised learning method that only leverages few hundreds of paired data and extra unpaired data for TTS and ASR. Our method consists of the following components: (1) a denoising auto-encoder, which reconstructs speech and text sequences respectively to develop the capability of language modeling both in speech and text domain; (2) dual transformation, where the TTS model transforms the text $y$ into speech $\hat{x}$, and the ASR model leverages the transformed pair $(\hat{x},y)$ for training, and vice versa, to boost the accuracy of the two tasks; (3) bidirectional sequence modeling, which addresses error propagation especially in the long speech and text sequence when training with few paired data; (4) a unified model structure, which combines all the above components for TTS and ASR based on Transformer model. Our method achieves 99.84% in terms of word level intelligible rate and 2.68 MOS for TTS, and 11.7% PER for ASR on LJSpeech dataset, by leveraging only 200 paired speech and text data (about 20 minutes audio), together with extra unpaired speech and text data.

Unsupervised Speech Recognition with N-Skipgram and Positional Unigram Matching

A Theory of Unsupervised Speech Recognition

Towards Unsupervised Speech Recognition Without Pronunciation Models

Analyzing the Robustness of Unsupervised Speech Recognition

Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation

Towards Unsupervised Automatic Speech Recognition Trained by Unaligned Speech and Text only

Learning Phone Recognition from Unpaired Audio and Phone Sequences Based on Generative Adversarial Network

Completely Unsupervised Phoneme Recognition By A Generative Adversarial Network Harmonized With Iteratively Refined Hidden Markov Models

Bridging Speech and Textual Pre-trained Models with Unsupervised ASR.

Robust Spoken Language Understanding With Unsupervised Asr-Error Adaptation

Unsupervised ASR via Cross-Lingual Pseudo-Labeling

Endpoint Detection for Streaming End-to-End Multi-Talker ASR

From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings.

Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition

Progressive Joint Modeling in Unsupervised Single-Channel Overlapped Speech Recognition.

Almost Unsupervised Text to Speech and Automatic Speech Recognition

Grammar-Supervised End-to-End Speech Recognition with Part-of-Speech Tagging and Dependency Parsing

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

The NPU-ASLP System for The ISCSLP 2022 Magichub Code-Swiching ASR Challenge

Improving Speech Recognition Using GAN-Based Speech Synthesis and Contrastive Unspoken Text Selection

Utilizing Self-supervised Representations for MOS Prediction