Abstract:Speech emotion recognition (SER) is one of the most challenging and active research topics in data science due to its wide range of applications in human–computer interaction, computer games, mobile services and psychological assessment. In the past, several studies have employed handcrafted features to classify emotions and achieved good classification accuracy. However, such features degrade the classification accuracy in complex scenarios. Thus, recent studies employed deep learning models to automatically extract the local representation from given audio signals. Though, automated feature engineering overcomes the issues of handcrafted feature extraction approach. However, still there is a need to further improve the performance of reported techniques. This is because, in reported techniques, single-layer and two-layer convolutional neural networks (CNNs) were used and these architectures are not capable of learning optimal features from complex speech signals. Thus, to overcome this limitation, this study proposed a novel SER framework, which applies data augmentation methods before extracting seven informative feature sets from each utterance. The extracted feature vector is used as input to the 1D CNN for emotions recognition using the EMO-DB, RAVDESS and SAVEE databases. Moreover, this study also proposed a cross-corpus SER model using the all audio files of common emotions of aforementioned databases. The experimental results showed that our proposed SER framework outperformed existing SER frameworks. Specifically, the proposed SER framework obtained 96.7% accuracy for EMO-DB with all utterances in seven emotions, 90.6% RAVDESS with all utterances in eight emotions, 93.2% for SAVEE with all utterances in seven emotions and 93.3% for cross-corpus with 1930 utterances in six emotions. We believe that our proposed framework will bring significant contribute to SER domain.

Improved multi-lingual sentiment analysis and recognition using deep learning

Improved Speech Emotion Classification Using Deep Neural Network

Cross-Corpus Speech Emotion Recognition Based on Hybrid Neural Networks

Deep Learning and Machine Learning-Based Model for Conversational Sentiment Classification

Speech Emotion Recognition Based on Convolutional Neural Network with Attention-Based Bidirectional Long Short-Term Memory Network and Multi-Task Learning

Convolutional neural network-based cross-corpus speech emotion recognition with data augmentation and features fusion

XEmoAccent: Embracing Diversity in Cross-Accent Emotion Recognition Using Deep Learning

Speech Emotion Recognition with Early Visual Cross-modal Enhancement Using Spiking Neural Networks.

Deep Learning-Based Sentiment Analysis for Roman Urdu Text

A Novel Approach for Sentiment Analysis of a Low Resource Language Using Deep Learning Models

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

Deep Sentiment Analysis Using CNN-LSTM Architecture of English and Roman Urdu Text Shared in Social Media

Exploiting Deep Learning for Persian Sentiment Analysis

Semi-supervised cross-lingual speech emotion recognition

Sentiment Analysis Based on Urdu Reviews Using Hybrid Deep Learning Models

Deep Cross-Corpus Speech Emotion Recognition: Recent Advances and Perspectives

Speech Emotion Recognition Using Mel-Frequency Cepstral Coefficients & Convolutional Neural Networks

Deep Learning, Ensemble and Supervised Machine Learning for Arabic Speech Emotion Recognition

A Hybrid Time-Distributed Deep Neural Architecture for Speech Emotion Recognition