Abstract:Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often struggle with inadequate speaker representation accuracy and overfitting, particularly in limited reference speeches scenarios. To address these challenges, we propose an Agile Speaker Representation Reinforcement Learning strategy to enhance speaker similarity in speaker adaptation tasks. ASRRL is the first work to apply reinforcement learning to improve the modeling accuracy of speaker embeddings in speaker adaptation, addressing the challenge of decoupling voice content and timbre. Our approach introduces two action strategies tailored to different reference speeches scenarios. In the single-sentence scenario, a knowledge-oriented optimal routine searching RL method is employed to expedite the exploration and retrieval of refinement information on the fringe of speaker representations. In the few-sentence scenario, we utilize a dynamic RL method to adaptively fuse reference speeches, enhancing the robustness and accuracy of speaker modeling. To achieve optimal results in the target domain, a multi-scale fusion scoring mechanism based reward model that evaluates speaker similarity, speech quality, and intelligibility across three dimensions is proposed, ensuring that improvements in speaker similarity do not compromise speech quality or intelligibility. The experimental results on the LibriTTS and VCTK datasets within mainstream TTS frameworks demonstrate the extensibility and generalization capabilities of the proposed ASRRL method. The results indicate that the ASRRL method significantly outperforms traditional fine-tuning approaches, achieving higher speaker similarity and better overall speech quality with limited reference speeches.

Variable Frame Rate Acoustic Models Using Minimum Error Reinforcement Learning.

Combining Frame-Synchronous and Label-Synchronous Systems for Speech Recognition

Unit Selection Speech Synthesis Using Frame-Sized Speech Segments and Neural Network Based Acoustic Models

Minimum word error training for non-autoregressive Transformer-based code-switching ASR

Hidden Markov Acoustic Modeling with Bootstrap and Restructuring for Low-Resourced Languages

FRS: Adaptive Score for Improving Acoustic Source Classification From Noisy Signals

Combining Hybrid DNN-HMM ASR Systems with Attention-Based Models Using Lattice Rescoring

Performance Improvements of Probabilistic Transcript-adapted ASR with Recurrent Neural Network and Language-specific Constraints

High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM Model

Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models

Multi-Span Acoustic Modelling Using Raw Waveform Signals.

Reinforcement Learning-based Frame-level Bit Allocation for VVC

Advanced Recurrent Network-Based Hybrid Acoustic Models for Low Resource Speech Recognition

Sequence-to-Sequence ASR Optimization via Reinforcement Learning

Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking

Frame Stacking and Retaining for Recurrent Neural Network Acoustic Model

Sequential Editing for Lifelong Training of Speech Recognition Models

ON MODULAR TRAINING OF NEURAL ACOUSTICS-TO-WORD MODEL FOR LVCSR

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

Integrating Lattice-Free MMI into End-to-End Speech Recognition

AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost