Text-independent speaker recognition using LSTM-RNN and speech enhancement
Samia Abd El-Moneim,M. A. Nassar,Moawad I. Dessouky,Nabil A. Ismail,Adel S. El-Fishawy,Fathi E. Abd El-Samie
DOI: https://doi.org/10.1007/s11042-019-08293-7
IF: 2.577
2020-06-17
Multimedia Tools and Applications
Abstract:Speaker recognition revolution has lead to the inclusion of speaker recognition modules in several commercial products. Most published algorithms for speaker recognition focus on text-dependent speaker recognition. In contrast, text-independent speaker recognition is more advantageous as the client can talk freely to the system. In this paper, text-independent speaker recognition is considered in the presence of some degradation effects such as noise and reverberation. Mel-Frequency Cepstral Coefficients (MFCCs), spectrum and log-spectrum are used for feature extraction from the speech signals. These features are processed with the Long-Short Term Memory Recurrent Neural Network (LSTM-RNN) as a classification tool to complete the speaker recognition task. The network learns to recognize the speakers efficiently in a text-independent manner, when the recording circumstances are the same. The recognition rate reaches 95.33% using MFCCs, while it is increased to 98.7% when using spectrum or log-spectrum. However, the system has some challenges to recognize speakers from different recording environments. Hence, different speech enhancement techniques, such as spectral subtraction and wavelet denoising, are used to improve the recognition performance to some extent. The proposed approach shows superiority, when compared to the algorithm of R. Togneri and D. Pullella (2011).
computer science, information systems, theory & methods,engineering, electrical & electronic, software engineering