Abstract:Speech production is a complex phenomenon, wherein the brain orchestrates a sequence of processes involving thought processing, motor planning, and the execution of articulatory movements. However, this intricate execution of various processes is susceptible to influence and disruption by various neurodegenerative pathological speech disorders, such as Parkinsons' disease, resulting in dysarthria, apraxia, and other conditions. These disorders lead to pathological speech characterized by abnormal speech patterns and imprecise articulation. Diagnosing these speech disorders in clinical settings typically involves auditory perceptual tests, which are time-consuming, and the diagnosis can vary among clinicians based on their experiences, biases, and cognitive load during the diagnosis. Additionally, unlike neurotypical speakers, patients with speech pathologies or impairments are unable to access various virtual assistants such as Alexa, Siri, etc. To address these challenges, several automatic pathological speech detection (PSD) approaches have been proposed. These approaches aim to provide efficient and accurate detection of speech disorders, thereby facilitating timely intervention and support for individuals affected by these conditions. These approaches mainly vary in two aspects: the input representations utilized and the classifiers employed. Due to the limited availability of data, the performance of detection remains subpar. Self-supervised learning (SSL) embeddings, such as wav2vec2, and their multilingual versions, are being explored as a promising avenue to improve performance. These embeddings leverage self-supervised learning techniques to extract rich representations from audio data, thereby offering a potential solution to address the limitations posed by the scarcity of labeled data.

Editorial Editorial of Special Issue on Self-Supervised Learning for Speech and Audio Processing

Guest Editorial: AI for Computational Audition—sound and Music Processing

Self-Supervised Speech Representation Learning: A Review

Self-supervised Representation Learning for Speech Processing

On‐the‐Job Search and the Wage Distribution

Investigating Self-Supervised Learning for Speech Enhancement and Separation

Efficient Personalized Speech Enhancement through Self-Supervised Learning

On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification

Progressive Multi-scale Self-supervised Learning for Speech Recognition

AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Efficiency-oriented approaches for self-supervised speech representation learning

Semi-Supervised Spoken Language Understanding Via Self-Supervised Speech and Language Model Pretraining.

The adult respiratory distress syndrome. Definition and prognosis.

Selfsupervised learning for pathological speech detection

Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models

Self-supervised models of audio effectively explain human cortical responses to speech

Self-Supervised Learning for Biomedical Signal Processing: A Systematic Review on ECG and PPG Signals

Introduction to the Special Issue on Deep Learning for Multi-Modal Intelligence Across Speech, Language, Vision, and Heterogeneous Signals

Improving Self-Supervised Learning for Audio Representations by Feature Diversity and Decorrelation

An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition