MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Qian Yang,Jialong Zuo,Zhe Su,Ziyue Jiang,Mingze Li,Zhou Zhao,Feiyang Chen,Zhefeng Wang,Baoxing Huai

2024-07-19

Abstract:We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at <a class="link-external link-https" href="https://speechai-demo.github.io/MSceneSpeech/" rel="external noopener nofollow">this https URL</a>.

Audio and Speech Processing,Sound

What problem does this paper attempt to address?

The paper aims to address the following issues: 1. **Lack of high-quality multi-scene speech datasets**: Existing speech synthesis datasets have limitations in terms of style diversity and prosody diversity, making it difficult to meet the needs of natural and expressive speech synthesis. To this end, the researchers created the MSceneSpeech dataset, which includes high-quality audio recordings from various everyday life scenarios. 2. **Challenges in cross-speaker style transfer**: In the task of cross-speaker style transfer, existing methods often struggle to handle both style transfer and voice adaptation simultaneously. The paper proposes a powerful baseline model that can achieve multi-style synthesis and cross-speaker style transfer, and operate based on any reference audio, making the process more convenient and unrestricted. 3. **Prosody control and separation**: Traditional methods face difficulties in independently controlling timbre and other style attributes (such as prosody). This paper proposes a new architecture that better separates timbre and prosody by explicitly modeling different components of prosody (duration, pitch, and energy). In summary, the main contributions of the paper are the release of a new multi-scene speech dataset, MSceneSpeech, and the proposal of a baseline model that effectively handles style transfer and cross-speaker adaptation, promoting the development of more expressive speech synthesis technology.

MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

KeSpeech: an Open Source Speech Dataset of Mandarin and Its Eight Subdialects.

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech Model

AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset

Design of Speech Corpus for Mandarin Text to Speech

SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

An Expressive Mandarin Speech Corpus

The Mandarin Chinese speech database: a corpus of 18,820 auditory neutral nonsense sentences

MES-P: an Emotional Tonal Speech Dataset in Mandarin Chinese with Distal and Proximal Labels

MES-P: an Emotional Tonal Speech Dataset in Mandarin with Distal and Proximal Labels

TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models

Articulatory Phonetics Informed Controllable Expressive Speech Synthesis

MnTTS: An Open-Source Mongolian Text-to-Speech Synthesis Dataset and Accompanied Baseline

Toward Synthesizing Expressive Mandarin Speech

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database

StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

MASS: Multi-task anthropomorphic speech synthesis framework