CDSD: Chinese Dysarthria Speech Database

Mengyi Sun,Ming Gao,Xinchen Kang,Shiru Wang,Jun Du,Dengfeng Yao,Su-Jing Wang

2023-10-24

Abstract:We present the Chinese Dysarthria Speech Database (CDSD) as a valuable resource for dysarthria research. This database comprises speech data from 24 participants with dysarthria. Among these participants, one recorded an additional 10 hours of speech data, while each recorded one hour, resulting in 34 hours of speech material. To accommodate participants with varying cognitive levels, our text pool primarily consists of content from the AISHELL-1 dataset and speeches by primary and secondary school students. When participants read these texts, they must use a mobile device or the ZOOM F8n multi-track field recorder to record their speeches. In this paper, we elucidate the data collection and annotation processes and present an approach for establishing a baseline for dysarthric speech recognition. Furthermore, we conducted a speaker-dependent dysarthric speech recognition experiment using an additional 10 hours of speech data from one of our participants. Our research findings indicate that, through extensive data-driven model training, fine-tuning limited quantities of specific individual data yields commendable results in speaker-dependent dysarthric speech recognition. However, we observe significant variations in recognition results among different dysarthric speakers. These insights provide valuable reference points for speaker-dependent dysarthric speech recognition.

Sound,Audio and Speech Processing

What problem does this paper attempt to address?

### Problems the Paper Aims to Solve This paper aims to address several key issues in Chinese Dysarthric Speech Recognition (DSR): 1. **Insufficient Data Scale**: Existing Chinese dysarthric speech databases are relatively small, with each database containing less than 10 hours of speech data. This limits the ability to train comprehensive speech recognition models. In contrast, the CDSD database proposed in this paper has a total duration of 34 hours, providing a larger scale of data resources. 2. **Ecological Validity**: Most participants recorded their speech using mobile devices at home, an environment that is closer to real-world application scenarios. This enhances the ecological validity of the data and the robustness of the models. 3. **Supplementary Video Data**: Researchers obtained permission from 6 participants to simultaneously record video data. By analyzing lip movements, they further confirmed the speech information and provided diverse data types for future research. 4. **Speaker-Dependent Speech Recognition**: Even with the same type of dysarthria, different individuals may have significant variations in pronouncing the same words. Therefore, the concept of speaker-dependent speech recognition is introduced to address the challenges of speech recognition for dysarthric speakers. In summary, this paper is primarily dedicated to constructing a large-scale, high-quality Chinese dysarthric speech database and exploring speaker-dependent speech recognition methods to improve the performance of dysarthric speech recognition systems.

CDSD: Chinese Dysarthria Speech Database

Pvd: A New Pathological Voice Dataset For Intra-Speaker Recognition Research Interest

Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design

Recent Progress in the CUHK Dysarthric Speech Recognition System

The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines

AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection

An undergraduate Mandarin speech database for speaker recognition research

Community-Supported Shared Infrastructure in Support of Speech Accessibility

CNAMD Corpus: A Chinese Natural Audiovisual Multimodal Database of Conversations for Social Interactive Agents

Design, construction and evaluation of emotional multimodal pathological speech database

KeSpeech: an Open Source Speech Dataset of Mandarin and Its Eight Subdialects.

Speech Databases for Mental Disorders: A Systematic Review

Construction and Evaluation of Mandarin Multimodal Emotional Speech Database

An Introduction to Mandarin Disordered Voice Database

Detecting somatisation disorder via speech: introducing the Shenzhen Somatisation Speech Corpus

A CANTONESE ACCENT CHINESE SPEECH CORPUS

OC16-CE80: A Chinese-English Mixlingual Database and A Speech Recognition Baseline

Chinese Emotional Speech Audiometry Project (CESAP): Establishment and Validation of a New Material Set With Emotionally Neutral Disyllabic Words

AISHELL-4 - An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Cantonese AphasiaBank: An annotated database of spoken discourse and co-verbal gestures by healthy and language-impaired native Cantonese speakers