DOI: 10.1145/3833088 ISSN: 1551-6857

Boosting Representation Learning for High-Level Semantic Information in Facial Expression Recognition

Hai Min, Kaiyi Jia, Chunxiao Fan, Peixi Yu, Wei Jia, Yang Zhao

Facial expression recognition (FER) requires high-level semantic cues to interpret abstract emotions. However, existing FER approaches commonly suffer from sparse semantic descriptions and unwanted noise from redundant facial pixels. To tackle these issues, we propose a boosting representation learning framework to build a robust high-level semantic projection space connecting textual semantics and facial images. To mitigate semantic scarcity for facial expression recognition, we develop the Expression High-level Semantic Encoding (EHSE) module. Guided by FACS-based physiological priors, this module constructs dense semantic prototypes to provide fine-grained supervised information for model learning. To remove irrelevant background noise mixed with redundant facial visual content, the Expression Key-core Features Focusing (EKFF) module is built, which picks out discriminative facial cues, boosting the signal-to-noise ratio and semantic purity of extracted visual features. To eliminate ambiguous classification boundaries within the semantic projection space, the Expression High-Level Semantic Alignment (EHSA) module is designed, which is equipped with multi-dimensional constraints. Experiments on FER+, RAF-DB and AffectNet verify the superiority of our approach. Our method obtains state-of-the-art results on FER+ and RAF-DB, and achieves the optimal average accuracy over AffectNet's 7-class and 8-class tasks. These results prove our framework can construct a balanced, discriminative and semantically intact projection space for FER.

More from our Archive