POSTER: A Pyramid Cross-Fusion Transformer Network for Facial Expression Recognition

Ce Zheng,Matias Mendieta,Chen Chen
2023-08-14
Abstract:Facial expression recognition (FER) is an important task in computer vision, having practical applications in areas such as human-computer interaction, education, healthcare, and online monitoring. In this challenging FER task, there are three key issues especially prevalent: inter-class similarity, intra-class discrepancy, and scale sensitivity. While existing works typically address some of these issues, none have fully addressed all three challenges in a unified framework. In this paper, we propose a two-stream Pyramid crOss-fuSion TransformER network (POSTER), that aims to holistically solve all three issues. Specifically, we design a transformer-based cross-fusion method that enables effective collaboration of facial landmark features and image features to maximize proper attention to salient facial regions. Furthermore, POSTER employs a pyramid structure to promote scale invariance. Extensive experimental results demonstrate that our POSTER achieves new state-of-the-art results on RAF-DB (92.05%), FERPlus (91.62%), as well as AffectNet 7 class (67.31%) and 8 class (63.34%). The code is available at <a class="link-external link-https" href="https://github.com/zczcwh/POSTER" rel="external noopener nofollow">this https URL</a>.
Computer Vision and Pattern Recognition,Artificial Intelligence
What problem does this paper attempt to address?
The paper aims to address three key issues in Facial Expression Recognition (FER): inter-class similarity, intra-class variability, and scale sensitivity. Specifically: - **Inter-class similarity**: Images of different expression categories can be very similar, with only subtle changes (such as slight changes in the mouth), making it difficult for existing methods to distinguish these images. - **Intra-class variability**: Images within the same expression category can have significant differences, such as variations in skin color, gender, age, and background. - **Scale sensitivity**: Deep learning networks are usually not stable when processing images of different resolutions and qualities. Although previous work has partially addressed these issues, no unified approach can simultaneously tackle all three challenges. Therefore, this paper proposes a Two-Stream Pyramid Cross-Fusion Transformer Network (POSTER), which combines facial landmarks and image features to address the aforementioned issues and achieves state-of-the-art results on multiple datasets.