Abstract:Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often struggle with inadequate speaker representation accuracy and overfitting, particularly in limited reference speeches scenarios. To address these challenges, we propose an Agile Speaker Representation Reinforcement Learning strategy to enhance speaker similarity in speaker adaptation tasks. ASRRL is the first work to apply reinforcement learning to improve the modeling accuracy of speaker embeddings in speaker adaptation, addressing the challenge of decoupling voice content and timbre. Our approach introduces two action strategies tailored to different reference speeches scenarios. In the single-sentence scenario, a knowledge-oriented optimal routine searching RL method is employed to expedite the exploration and retrieval of refinement information on the fringe of speaker representations. In the few-sentence scenario, we utilize a dynamic RL method to adaptively fuse reference speeches, enhancing the robustness and accuracy of speaker modeling. To achieve optimal results in the target domain, a multi-scale fusion scoring mechanism based reward model that evaluates speaker similarity, speech quality, and intelligibility across three dimensions is proposed, ensuring that improvements in speaker similarity do not compromise speech quality or intelligibility. The experimental results on the LibriTTS and VCTK datasets within mainstream TTS frameworks demonstrate the extensibility and generalization capabilities of the proposed ASRRL method. The results indicate that the ASRRL method significantly outperforms traditional fine-tuning approaches, achieving higher speaker similarity and better overall speech quality with limited reference speeches.

Towards Automatic Data Augmentation for Disordered Speech Recognition

Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition

Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition

Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation

Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis

Adaptive data augmentation for mandarin automatic speech recognition

Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation

Auditory-Based Data Augmentation for End-to-End Automatic Speech Recognition

Data Augmentation for End-to-end Code-switching Speech Recognition

Data Augmentation with Locally-time Reversed Speech for Automatic Speech Recognition

Semantic Data Augmentation for End-to-End Mandarin Speech Recognition

Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation

ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR

SpecSwap: A Simple Data Augmentation Method for End-to-End Speech Recognition

Automatically Learning Data Augmentation Policies for Dialogue Tasks

PhasePerturbation: Speech Data Augmentation via Phase Perturbation for Automatic Speech Recognition

Data Augmentation for Arabic Speech Recognition Based on End-to-End Deep Learning

Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

On the effectiveness of enrollment speech augmentation for Target Speaker Extraction

Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis