Abstract:There has been a growing interest in multimodal sentiment analysis and emotion recognition in recent years due to its wide range of practical applications. Multiple modalities allow for the integration of complementary information, improving the accuracy and precision of sentiment and emotion recognition tasks. However, working with multiple modalities presents several challenges, including handling data source heterogeneity, fusing information, aligning and synchronizing modalities, and designing effective feature extraction techniques that capture discriminative information from each modality. This paper introduces a novel framework called "Attention-based Multimodal Sentiment Analysis and Emotion Recognition (AMSAER)" to address these challenges. This framework leverages intra-modality discriminative features and inter-modality correlations in visual, audio, and textual modalities. It incorporates an attention mechanism to facilitate sentiment and emotion classification based on visual, textual, and acoustic inputs by emphasizing relevant aspects of the task. The proposed approach employs separate models for each modality to automatically extract discriminative semantic words, image regions, and audio features. A deep hierarchical model is then developed, incorporating intermediate fusion to learn hierarchical correlations between the modalities at bimodal and trimodal levels. Finally, the framework combines four distinct models through decision-level fusion to enable multimodal sentiment analysis and emotion recognition. The effectiveness of the proposed framework is demonstrated through extensive experiments conducted on the publicly available Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. The results confirm a notable performance improvement compared to state-of-the-art methods, attaining 85% and 93% accuracy for sentiment analysis and emotion classification, respectively. Additionally, when considering class-wise accuracy, the results indicate that the "angry" emotion and "positive" sentiment are classified more effectively than the other emotions and sentiments, achieving 96.80% and 93.14% accuracy, respectively.

Multimodal modelling of human emotion using sound, image and text fusion

Emotion Recognition in Videos via Fusing Multimodal Features.

Investigating Multisensory Integration in Emotion Recognition Through Bio-Inspired Computational Models

Multimodal Emotional Classification Based on Meaningful Learning

Multimodal emotion recognition using cross modal audio-video fusion with attention and deep metric learning

Multimodal emotion recognition model via hybrid model with improved feature level fusion on facial and EEG feature set

Multimodal Emotion Recognition Using Different Fusion Techniques

Multimodal Emotion Recognition Based on Cascaded Multichannel and Hierarchical Fusion

Deep learning based multimodal emotion recognition using model-level fusion of audio–visual modalities

A multimodal emotion recognition model integrating speech, video and MoCAP

Multimodal Emotion Recognition using Transfer Learning from Speaker Recognition and BERT-based models

Multimodal Emotion Recognition by Extracting Common and Modality-Specific Information.

Multimodal Emotion Recognition Based on Feature Fusion.

Multimodal Emotion Detection via Attention-Based Fusion of Extracted Facial and Speech Features

Multi-head attention fusion networks for multi-modal speech emotion recognition

Multimodal Utterance-level Affect Analysis using Visual, Audio and Text Features

Multimodal emotion recognition based on a fusion of audiovisual information with temporal dynamics

Multimodal Speech Emotion Recognition Using Audio and Text

Attention-based multimodal sentiment analysis and emotion recognition using deep neural networks

Multimodal interaction enhanced representation learning for video emotion recognition

Emotion Recognition Model Based on Multimodal Decision Fusion