Abstract:Speech emotion recognition (SER) is one of the most challenging and active research topics in data science due to its wide range of applications in human–computer interaction, computer games, mobile services and psychological assessment. In the past, several studies have employed handcrafted features to classify emotions and achieved good classification accuracy. However, such features degrade the classification accuracy in complex scenarios. Thus, recent studies employed deep learning models to automatically extract the local representation from given audio signals. Though, automated feature engineering overcomes the issues of handcrafted feature extraction approach. However, still there is a need to further improve the performance of reported techniques. This is because, in reported techniques, single-layer and two-layer convolutional neural networks (CNNs) were used and these architectures are not capable of learning optimal features from complex speech signals. Thus, to overcome this limitation, this study proposed a novel SER framework, which applies data augmentation methods before extracting seven informative feature sets from each utterance. The extracted feature vector is used as input to the 1D CNN for emotions recognition using the EMO-DB, RAVDESS and SAVEE databases. Moreover, this study also proposed a cross-corpus SER model using the all audio files of common emotions of aforementioned databases. The experimental results showed that our proposed SER framework outperformed existing SER frameworks. Specifically, the proposed SER framework obtained 96.7% accuracy for EMO-DB with all utterances in seven emotions, 90.6% RAVDESS with all utterances in eight emotions, 93.2% for SAVEE with all utterances in seven emotions and 93.3% for cross-corpus with 1930 utterances in six emotions. We believe that our proposed framework will bring significant contribute to SER domain.

Feature selection enhancement and feature space visualization for speech-based emotion recognition

Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion Recognition

Deep Spectrum Feature Representations for Speech Emotion Recognition

A feature selection model for speech emotion recognition using clustering-based population generation with hybrid of equilibrium optimizer and atom search optimization algorithm

An Improved MSER using Grid Search based PCA and Ensemble Voting Technique

Speaker-independent Speech Emotion Recognition Based on Random Forest Feature Selection Algorithm

Visual-Audio Emotion Recognition Based on Multi-Task and Ensemble Learning with Multiple Features

Speech Emotion Recognition Based on Syllable-Level Feature Extraction

Real-time Speech Emotion Recognition Based on Syllable-Level Feature Extraction

Survey on Discriminative Feature Selection for Speech Emotion Recognition

Speech Emotion Recognition Based on Formant Characteristics Feature Extraction and Phoneme Type Convergence.

Fusion of PCA and ICA in Statistical Subset Analysis for Speech Emotion Recognition

Feature extraction algorithms to improve the speech emotion recognition rate

Enhancing speech emotion recognition through deep learning and handcrafted feature fusion

Emotion Recognition From Noisy Speech

Exploring Language-Independent Emotional Acoustic Features via Feature Selection

Study on Feature Subspace of Archetypal Emotions for Speech Emotion Recognition

Improved Speech Emotion Classification Using Deep Neural Network

Comparative Performance Analysis of Metaheuristic Feature Selection Methods for Speech Emotion Recognition

Convolutional neural network-based cross-corpus speech emotion recognition with data augmentation and features fusion