Abstract:Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. Sources of heterogeneity commonly found in normal speech including accent or gender, when further compounded with the variability over age and speech pathology severity level, create large diversity among speakers. To this end, speaker adaptation techniques play a key role in personalization of ASR systems for such users. Motivated by the spectro-temporal level differences between dysarthric, elderly and normal speech that systematically manifest in articulatory imprecision, decreased volume and clarity, slower speaking rates and increased dysfluencies, novel spectrotemporal subspace basis deep embedding features derived using SVD speech spectrum decomposition are proposed in this paper to facilitate auxiliary feature based speaker adaptation of state-of-the-art hybrid DNN/TDNN and end-to-end Conformer speech recognition systems. Experiments were conducted on four tasks: the English UASpeech and TORGO dysarthric speech corpora; the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The proposed spectro-temporal deep feature adapted systems outperformed baseline i-Vector and xVector adaptation by up to 2.63% absolute (8.63% relative) reduction in word error rate (WER). Consistent performance improvements were retained after model based speaker adaptation using learning hidden unit contributions (LHUC) was further applied. The best speaker adapted system using the proposed spectral basis embedding features produced the lowest published WER of 25.05% on the UASpeech test set of 16 dysarthric speakers.

Multi-Channel Feature Adaptation for Robust Speech Recognition

Noise Robust Speech Recognition Using Multi-Channel Based Channel Selection And ChannelWeighting.

A Feature Integration Network for Multi-Channel Speech Enhancement

Automatic channel selection and spatial feature integration for multi-channel speech recognition across various array topologies

Channel selection using neural network posterior probability for speech recognition with distributed microphone arrays in everyday environments

On Design of Robust Deep Models for CHiME-4 Multi-Channel Speech Recognition with Multiple Configurations of Array Microphones

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention

Robust speech recognition using beamforming with adaptive microphone gains and multichannel noise reduction

Phoneme Dependent Speaker Embedding And Model Factorization For Multi-Speaker Speech Synthesis And Adaptation

Factorised Speaker-environment Adaptive Training of Conformer Speech Recognition Systems

Adapting Multi-Lingual ASR Models for Handling Multiple Talkers

Unsupervised Adaptation with Domain Separation Networks for Robust Speech Recognition

Rapid Adaptation For Deep Neural Networks Through Multi-Task Learning

Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition

3-D Feature and Acoustic Modeling for Far-Field Speech Recognition

Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement

Multi-feature Combination for Speaker Recognition

End-to-end Multichannel Speaker-Attributed ASR: Speaker Guided Decoder and Input Feature Analysis

Deep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognition

Improved Frequency Modulation Features for Multichannel Distant Speech Recognition