Abstract:Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthesis models currently rely on annotated audio data, but it is crucial to scale them to self-supervised datasets in order to effectively capture the wide range of acoustic variations present in human voice, including speaker identity, emotion, and prosody. In this work, we propose Make-A-Voice, a unified framework for synthesizing and manipulating voice signals from discrete representations. Make-A-Voice leverages a "coarse-to-fine" approach to model the human voice, which involves three stages: 1) semantic stage: model high-level transformation between linguistic content and self-supervised semantic tokens, 2) acoustic stage: introduce varying control signals as acoustic conditions for semantic-to-acoustic modeling, and 3) generation stage: synthesize high-fidelity waveforms from acoustic tokens. Make-A-Voice offers notable benefits as a unified voice synthesis framework: 1) Data scalability: the major backbone (i.e., acoustic and generation stage) does not require any annotations, and thus the training data could be scaled up. 2) Controllability and conditioning flexibility: we investigate different conditioning mechanisms and effectively handle three voice synthesis applications, including text-to-speech (TTS), voice conversion (VC), and singing voice synthesis (SVS) by re-synthesizing the discrete voice representations with prompt guidance. Experimental results demonstrate that Make-A-Voice exhibits superior audio quality and style similarity compared with competitive baseline models. Audio samples are available at <a class="link-external link-https" href="https://Make-A-Voice.github.io" rel="external noopener nofollow">this https URL</a>

High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

AudioVSR: Enhancing Video Speech Recognition with Audio Data

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

Supervisory Data Alignment for Text-Independent Voice Conversion

Video-to-Audio Generation with Hidden Alignment

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

Make-A-Voice: Unified Voice Synthesis With Discrete Representation

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

Almost Unsupervised Text to Speech and Automatic Speech Recognition

Mutual Alignment between Audiovisual Features for End-to-End Audiovisual Speech Recognition

Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

Audio-visual speech synthesis using vision transformer–enhanced autoencoders with ensemble of loss functions

Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation

Synth-AC: Enhancing Audio Captioning with Synthetic Supervision

Unified Cross-Modal Attention: Robust Audio-Visual Speech Recognition and Beyond

When Automatic Voice Disguise Meets Automatic Speaker Verification

Aligning Audio-Visual Joint Representations with an Agentic Workflow