Abstract:There has been a growing interest in multimodal sentiment analysis and emotion recognition in recent years due to its wide range of practical applications. Multiple modalities allow for the integration of complementary information, improving the accuracy and precision of sentiment and emotion recognition tasks. However, working with multiple modalities presents several challenges, including handling data source heterogeneity, fusing information, aligning and synchronizing modalities, and designing effective feature extraction techniques that capture discriminative information from each modality. This paper introduces a novel framework called "Attention-based Multimodal Sentiment Analysis and Emotion Recognition (AMSAER)" to address these challenges. This framework leverages intra-modality discriminative features and inter-modality correlations in visual, audio, and textual modalities. It incorporates an attention mechanism to facilitate sentiment and emotion classification based on visual, textual, and acoustic inputs by emphasizing relevant aspects of the task. The proposed approach employs separate models for each modality to automatically extract discriminative semantic words, image regions, and audio features. A deep hierarchical model is then developed, incorporating intermediate fusion to learn hierarchical correlations between the modalities at bimodal and trimodal levels. Finally, the framework combines four distinct models through decision-level fusion to enable multimodal sentiment analysis and emotion recognition. The effectiveness of the proposed framework is demonstrated through extensive experiments conducted on the publicly available Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. The results confirm a notable performance improvement compared to state-of-the-art methods, attaining 85% and 93% accuracy for sentiment analysis and emotion classification, respectively. Additionally, when considering class-wise accuracy, the results indicate that the "angry" emotion and "positive" sentiment are classified more effectively than the other emotions and sentiments, achieving 96.80% and 93.14% accuracy, respectively.

Semantic Alignment Network for Multi-modal Emotion Recognition

Bridging the Emotional Semantic Gap via Multimodal Relevance Estimation

Investigating Multisensory Integration in Emotion Recognition Through Bio-Inspired Computational Models

AMSA: Adaptive Multimodal Learning for Sentiment Analysis

Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition

Research on cross-modal emotion recognition based on multi-layer semantic fusion

A Multi-Level Alignment and Cross-Modal Unified Semantic Graph Refinement Network for Conversational Emotion Recognition

Semantic Alignment for Multimodal Large Language Models

Multi-modal fusion network with complementarity and importance for emotion recognition

MSAF: Multimodal Split Attention Fusion

AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations

Multi-head attention fusion networks for multi-modal speech emotion recognition

Semantic-specific multimodal relation learning for sentiment analysis

Multi-layer cross-modality attention fusion network for multimodal sentiment analysis

A multimodal shared network with a cross-modal distribution constraint for continuous emotion recognition

Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment

Multi-modal Attention for Speech Emotion Recognition

SAM: Modeling Scene, Object and Action with Semantics Attention Modules for Video Recognition

Attention-based multimodal sentiment analysis and emotion recognition using deep neural networks

Learning Semantic Alignment Using Global Features and Multi-scale Confidence

Multi-Modal Sentiment Analysis Based on Image and Text Fusion Based on Cross-Attention Mechanism