Abstract:Considerable attention has been paid to physiological signal-based emotion recognition in the field of affective computing. For reliability and user-friendly acquisition, electrodermal activity (EDA) has a great advantage in practical applications. However, EDA-based emotion recognition with large-scale subjects is still a tough problem. The traditional well-designed classifiers with hand-crafted features produce poorer results because of their limited representation abilities. And the deep learning models with auto feature extraction suffer the overfitting drop-off because of large-scale individual differences. Since music has a strong correlation with human emotion, static music can be involved as the external benchmark to constrain various dynamic EDA signals. In this article, we make an attempt by fusing the subject’s individual EDA features and the external evoked music features. And we propose an end-to-end multimodal framework, the one-dimensional residual temporal and channel attention network (RTCAN-1D). For EDA features, the channel-temporal attention mechanism for EDA-based emotion recognition is first involved in mine the temporal and channel-wise dynamic and steady features. The comparisons with single EDA-based SOTA models on DEAP and AMIGOS datasets prove the effectiveness of RTCAN-1D to mine EDA features. For music features, we simply process the music signal with the open-source toolkit openSMILE to obtain external feature vectors. We conducted systematic and extensive evaluations. The experiments on the current largest music emotion dataset PMEmo validate that the fusion of EDA and music is a reliable and efficient solution for large-scale emotion recognition.

Frequency Embedded Regularization Network for Continuous Music Emotion Recognition

A Efficient Multimodal Framework for Large Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals

Deep Spectrum Feature Representations for Speech Emotion Recognition

ADFF: Attention Based Deep Feature Fusion Approach for Music Emotion Recognition

MFDR: Multiple-stage Fusion and Dynamically Refined Network for Multimodal Emotion Recognition

Semi-Supervised Self-Learning Enhanced Music Emotion Recognition

Music-induced emotion flow modeling by ENMI Network

Music emotion recognition using deep convolutional neural networks

FFA-BiGRU: Attention-Based Spatial-Temporal Feature Extraction Model for Music Emotion Classification

A Multimodal Framework for Large-Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals

Symbolic & Acoustic: Multi-domain Music Emotion Modeling for Instrumental Music

Fusion of EEG and Musical Features in Continuous Music-emotion Recognition

Feature Selection Approaches for Optimising Music Emotion Recognition Methods

The PMEmo Dataset for Music Emotion Recognition

Music Emotion Recognition Based on a Neural Network with an Inception-GRU Residual Structure

Music Emotion Prediction Using Recurrent Neural Networks

Music emotion recognition based on temporal convolutional attention network using EEG

Research on Music Emotional Expression Based on Reinforcement Learning and Multimodal Information

Continuous Emotion Recognition during Music Listening Using EEG Signals: A Fuzzy Parallel Cascades Model

IIOF: Intra- and Inter-feature orthogonal fusion of local and global features for music emotion recognition

A Comparison Study of Deep Learning Methodologies for Music Emotion Recognition