Abstract:In-the-wild Dynamic facial expression recognition (DFER) encounters a significant challenge in recognizing emotion-related expressions, which are often temporally and spatially diluted by emotion-irrelevant expressions and global context respectively. Most of the prior DFER methods model tightly coupled spatiotemporal representations which may incorporate weakly relevant features, leading to information redundancy and emotion-irrelevant context bias. Several DFER methods have highlighted the significance of dynamic information, but utilize explicit manners to extract dynamic features with overly strong prior knowledge. In this paper, we propose a novel Implicit Facial Dynamics Disentanglement framework (IFDD). Through expanding wavelet lifting scheme to fully learnable framework, IFDD disentangles emotion-related dynamic information from emotion-irrelevant global context in an implicit manner, i.e., without exploit operations and external guidance. The disentanglement process of IFDD contains two stages, i.e., Inter-frame Static-dynamic Splitting Module (ISSM) for rough disentanglement estimation and Lifting-based Aggregation-Disentanglement Module (LADM) for further refinement. Specifically, ISSM explores inter-frame correlation to generate content-aware splitting indexes on-the-fly. We preliminarily utilize these indexes to split frame features into two groups, one with greater global similarity, and the other with more unique dynamic features. Subsequently, LADM first aggregates these two groups of features to obtain fine-grained global context features by an updater, and then disentangles emotion-related facial dynamic features from the global context by a predictor. Extensive experiments on in-the-wild datasets have demonstrated that IFDD outperforms prior supervised DFER methods with higher recognition accuracy and comparable efficiency.

Hidden Markov Model Decision Forest for Dynamic Facial Expression Recognition

MFDR: Multiple-stage Fusion and Dynamically Refined Network for Multimodal Emotion Recognition

Real-Time facial expression recognition system based on HMM and feature point localization

Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition

Early Facial Expression Recognition Using Hidden Markov Models

Speech Emotion Recognition Based on Linear Discriminant Analysis and Support Vector Machine Decision Tree

Combining 2D Gabor and Local Binary Pattern for Facial Expression Recognition Using Extreme Learning Machine

Multi-Instance Hidden Markov Model For Facial Expression Recognition

Adaptive key-frame selection-based facial expression recognition via multi-cue dynamic features hybrid fusion

The heterogeneous ensemble of deep forest and deep neural networks for micro-expressions recognition

Spatio-temporal deep forest for emotion recognition based on facial electromyography signals

Sparse Modified Marginal Fisher Analysis for Facial Expression Recognition

A Spatio-Temporal Integrated Model Based on Local and Global Features for Video Expression Recognition

Freq-HD: an Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for In-the-wild Facial Expression Recognition in Videos

Expression recognition from video using a coupled hidden Markov model

Dynamic Facial Expression Recognition Using Boosted Component-Based Spatiotemporal Features and Multi-classifier Fusion

Facial Expression Recognition from Image Sequences Using Twofold Random Forest Classifier

Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild

Seeing Through the Mask: Recognition of Genuine Emotion Through Masked Facial Expression

The investigation of traditional models and machine learning models in dynamic facial expression recognition

Facial Expression Recognition Based on Multi‐regional D–S Evidences Theory Fusion