Abstract:Human affective behavior analysis focuses on analyzing human expressions or other behaviors to enhance the understanding of human psychology. The CVPR 2023 Competition on Affective Behavior Analysis in-the-wild (ABAW) is dedicated to providing high-quality and large-scale Aff-wild2 for the recognition of commonly used emotion representations, such as Action Units (AU), basic expression categories(EXPR), and Valence-Arousal (VA). The competition is committed to making significant strides in improving the accuracy and practicality of affective analysis research in real-world scenarios. In this paper, we introduce our submission to the CVPR 2023: ABAW5. Our approach involves several key components. First, we utilize the visual information from a Masked Autoencoder(MAE) model that has been pre-trained on a large-scale face image dataset in a self-supervised manner. Next, we finetune the MAE encoder on the image frames from the Aff-wild2 for AU, EXPR and VA tasks, which can be regarded as a static and uni-modal training. Additionally, we leverage the multi-modal and temporal information from the videos and implement a transformer-based framework to fuse the multi-modal features. Our approach achieves impressive results in the ABAW5 competition, with an average F1 score of 55.49\% and 41.21\% in the AU and EXPR tracks, respectively, and an average CCC of 0.6372 in the VA track. Our approach ranks first in the EXPR and AU tracks, and second in the VA track. Extensive quantitative experiments and ablation studies demonstrate the effectiveness of our proposed method.

Facial Affect Recognition based on Transformer Encoder and Audiovisual Fusion for the ABAW5 Challenge

Facial Affect Recognition based on Multi Architecture Encoder and Feature Fusion for the ABAW7 Challenge

Facial Affective Behavior Analysis Method for 5th ABAW Competition

Facial Expression Recognition Based on Multi-modal Features for Videos in the Wild

An Effective Ensemble Learning Framework for Affective Behaviour Analysis

Transformer-based Multimodal Information Fusion for Facial Expression Analysis

Multi-modal Facial Action Unit Detection with Large Pre-trained Models for the 5th Competition on Affective Behavior Analysis in-the-wild

A Efficient Multimodal Framework for Large Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals

Affective Behaviour Analysis via Integrating Multi-Modal Knowledge

Multimodal Feature Extraction and Fusion for Emotional Reaction Intensity Estimation and Expression Classification in Videos with Transformers

Multi-modal Facial Affective Analysis based on Masked Autoencoder

Facial Affect Analysis: Learning from Synthetic Data & Multi-Task Learning Challenges

ABAW: Valence-Arousal Estimation, Expression Recognition, Action Unit Detection & Emotional Reaction Intensity Estimation Challenges

Multi-modal Expression Recognition with Ensemble Method

A Unified Approach to Facial Affect Analysis: the MAE-Face Visual Representation.

Spatial-temporal Transformer for Affective Behavior Analysis

Multi-Task Learning for Emotion Descriptors Estimation at the fourth ABAW Challenge

ABAW : Facial Expression Recognition in the wild

Visual-Audio Emotion Recognition Based on Multi-Task and Ensemble Learning with Multiple Features

Mutilmodal Feature Extraction and Attention-based Fusion for Emotion Estimation in Videos

Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation