Abstract:Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the emotional style modeling and speaker identity conversion in **Expressive Voice Conversion (EVC)**. Specifically, the paper focuses on how to simultaneously convert the speaker's identity and emotional style while keeping the language content unchanged, especially the challenges faced when dealing with the emotional voices of arbitrary speakers. Traditional methods mainly focus on neutral voices, ignoring the prosodic changes in different emotional states, and relying on vocoders for voice reconstruction, which makes the quality of the synthesized voice highly dependent on the performance of the vocoder. In addition, emotional styles contain speaker - independent and speaker - dependent features, and these features have not been fully explored in previous studies. To address the above challenges, this paper proposes an end - to - end emotional voice conversion framework (DEVC) based on the Conditional Denoising Diffusion Probabilistic Model (DDPM). This framework uses speech units extracted by self - supervised speech models as content conditions, and combines deep features extracted from speech emotion recognition and speaker verification systems to model emotional styles and speaker identities. Through this method, DEVC can more effectively capture and convert the emotional characteristics of speakers, achieving high - quality any - to - any emotional voice conversion. The main contributions of the paper include: - Proposing an end - to - end emotional voice conversion framework that does not require large - scale training data and manual annotation; - Discovering that the speaker embeddings extracted by the speaker verification model pre - trained with neutral data can effectively capture speaker - dependent emotional cues, thereby enhancing the effect of emotional voice conversion; - The framework demonstrates the flexibility of identity conversion for known and unknown emotional speakers, achieving any - to - any emotional voice conversion. Through objective and subjective evaluations, the research has proven the effectiveness of the proposed framework.

Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

EXPRESSIVE VOICE CONVERSION: A JOINT FRAMEWORK FOR SPEAKER IDENTITY AND EMOTIONAL STYLE TRANSFER

Expressive-VC: Highly Expressive Voice Conversion with Attention Fusion of Bottleneck and Perturbation Features

PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice Conversion

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

Toward Any-to-Any Emotion Voice Conversion using Disentangled Diffusion Framework

HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts

DurFlex-EVC: Duration-Flexible Emotional Voice Conversion with Parallel Generation

Highly Controllable Diffusion-based Any-to-Any Voice Conversion Model with Frame-level Prosody Feature

Multi-Target Emotional Voice Conversion With Neural Vocoders

Decoupling Speaker-Independent Emotions for Voice Conversion Via Source-Filter Networks

Emotional Voice Conversion With Cycle-consistent Adversarial Network

DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion

Voice Conversion with Denoising Diffusion Probabilistic GAN Models

Towards General-Purpose Text-Instruction-Guided Voice Conversion

Nonparallel Emotional Voice Conversion For Unseen Speaker-Emotion Pairs Using Dual Domain Adversarial Network & Virtual Domain Pairing

Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training

Mixed-EVC: Mixed Emotion Synthesis and Control in Voice Conversion

An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning

Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion

CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion