Abstract:Goal: As an essential human-machine interactive task, emotion recognition has become an emerging area over the decades. Although previous attempts to classify emotions have achieved high performance, several challenges remain open: 1) How to effectively recognize emotions using different modalities remains challenging. 2) Due to the increasing amount of computing power required for deep learning, how to provide real-time detection and improve the robustness of deep neural networks is important. Method: In this paper, we propose a deep learning-based multimodal emotion recognition (MER) called Deep-Emotion, which can adaptively integrate the most discriminating features from facial expressions, speech, and electroencephalogram (EEG) to improve the performance of the MER. Specifically, the proposed Deep-Emotion framework consists of three branches, i.e., the facial branch, speech branch, and EEG branch. Correspondingly, the facial branch uses the improved GhostNet neural network proposed in this paper for feature extraction, which effectively alleviates the overfitting phenomenon in the training process and improves the classification accuracy compared with the original GhostNet network. For work on the speech branch, this paper proposes a lightweight fully convolutional neural network (LFCNN) for the efficient extraction of speech emotion features. Regarding the study of EEG branches, we proposed a tree-like LSTM (tLSTM) model capable of fusing multi-stage features for EEG emotion feature extraction. Finally, we adopted the strategy of decision-level fusion to integrate the recognition results of the above three modes, resulting in more comprehensive and accurate performance. Result and Conclusions: Extensive experiments on the CK+, EMO-DB, and MAHNOB-HCI datasets have demonstrated the advanced nature of the Deep-Emotion method proposed in this paper, as well as the feasibility and superiority of the MER approach.

Student's Emotion Recognition using Multimodality and Deep Learning

Multimodal Emotional Classification Based on Meaningful Learning

Multimodal Emotion Recognition Using Different Fusion Techniques

Multimodal Emotion Recognition and State Analysis of Classroom Video and Audio Based on Deep Neural Network

Classroom Student Emotions Classification from Facial Expressions and Speech Signals using Deep Learning

Multimodal Emotion Recognition System Using Machine Learning Classifier

Multimodal Emotion Recognition by Combining Physiological Signals and Facial Expressions: a Preliminary Study.

Emotion recognition using multimodal deep learning in multiple psychophysiological signals and video

Real-time emotional health detection using fine-tuned transfer networks with multimodal fusion

Deep learning based multimodal emotion recognition using model-level fusion of audio–visual modalities

A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product

Emotion recognition framework using multiple modalities for an effective human–computer interaction

Multimodal Physiological-Based Emotion Recognition

Multimodal emotion recognition model via hybrid model with improved feature level fusion on facial and EEG feature set

Multimodal Emotion Recognition based on Facial Expressions, Speech, and EEG

Emotion Recognition System via Facial Expressions and Speech Using Machine Learning and Deep Learning Techniques

Novel multimodal emotion detection method using Electroencephalogram and Electrocardiogram signals

Text-Based Emotion Recognition Using Deep Learning Approach

Emotion Recognition for Challenged People Facial Appearance in Social using Neural Network

Speech emotion recognition using multimodal feature fusion with machine learning approach

Multimodal Emotion Recognition using Transfer Learning from Speaker Recognition and BERT-based models