Abstract:Convolutional neural networks (CNNs) have made significant progress in the field of facial expression recognition (FER). However, due to challenges such as occlusion, lighting variations, and changes in head pose, facial expression recognition in real-world environments remains highly challenging. At the same time, methods solely based on CNN heavily rely on local spatial features, lack global information, and struggle to balance the relationship between computational complexity and recognition accuracy. Consequently, the CNN-based models still fall short in their ability to address FER adequately. To address these issues, we propose a lightweight facial expression recognition method based on a hybrid vision transformer. This method captures multi-scale facial features through an improved attention module, achieving richer feature integration, enhancing the network's perception of key facial expression regions, and improving feature extraction capabilities. Additionally, to further enhance the model's performance, we have designed the patch dropping (PD) module. This module aims to emulate the attention allocation mechanism of the human visual system for local features, guiding the network to focus on the most discriminative features, reducing the influence of irrelevant features, and intuitively lowering computational costs. Extensive experiments demonstrate that our approach significantly outperforms other methods, achieving an accuracy of 86.51% on RAF-DB and nearly 70% on FER2013, with a model size of only 3.64 MB. These results demonstrate that our method provides a new perspective for the field of facial expression recognition.

Attend to Where and When: Cascaded Attention Network for Facial Expression Recognition

A Channel-Wise Spatial-Temporal Aggregation Network for Action Recognition

A Cascade Attention Based Facial Expression Recognition Network by Fusing Multi-Scale Spatio-Temporal Features

CANet: Comprehensive Attention Network for video-based action recognition

Attention mechanism-based CNN for facial expression recognition

SAANet: Siamese Action-Units Attention Network for Improving Dynamic Facial Expression Recognition

Distract Your Attention: Multi-Head Cross Attention Network for Facial Expression Recognition

An Efficient Channel Attention CNN for Facial Expression Recognition

Hybrid Attention-Aware Learning Network for Facial Expression Recognition in the Wild

ATTENTION BASED CONVOLUTIONAL NEURAL NETWORK FOR FACIAL EXPRESSION RECOGNITION

Coarse-to-Fine Cascaded Networks with Smooth Predicting for Video Facial Expression Recognition

Spatiotemporal Convolutional Neural Network with Convolutional Block Attention Module for Micro-Expression Recognition

TriCAFFNet: A Tri-Cross-Attention Transformer with a Multi-Feature Fusion Network for Facial Expression Recognition

Enhanced Hybrid Vision Transformer with Multi-Scale Feature Integration and Patch Dropping for Facial Expression Recognition

Multi-Stream Facial Adaptive Network for Expression Recognition from a Single Image

Facial Expression Recognition Based on Deep Evolutional Spatial-Temporal Networks

A multi-scale multi-attention network for dynamic facial expression recognition

Adaptive multilayer perceptual attention network for facial expression recognition

HiCAN: Hierarchical Convolutional Attention Network for Sequence Modeling.

PASTFNet: a paralleled attention spatio-temporal fusion network for micro-expression recognition