ATSLA: Attention-Driven Temporal–Spatial Learning Architecture for Dynamic Facial Expression Recognition
Hamza Ghulam Nabi, Kowovi Comivi Alowonou, Ji-Hyeong HanDynamic facial expression recognition (DFER) is a crucial field in computer vision that aims to automatically detect and analyze human emotions from video sequences. Despite recent advances, current DFER methods face critical limitations. Existing approaches struggle to effectively integrate global facial features with detailed local information, temporal modeling techniques inadequately capture subtle expression changes, and most methods are vulnerable to noisy annotations from crowdsourced datasets. This study proposes a novel framework, an attention-driven temporal–spatial learning architecture for DFER (ATSLA-DFER), designed to address these critical issues in DFER tasks. The proposed approach leverages a multi-component architecture that effectively combines spatial and temporal learning mechanisms. It first incorporates a dual-stream feature extraction process utilizing pre-trained IR50 and MobileFaceNet backbones for micro- and macro-spatial feature processing. The global local attention module (GLAM) further enhances the macro features for the feature-enriching process. Moreover, a cross-adaptive feature fusion (CAFF) is proposed for effective multi-scale feature integration, and a custom-designed temporal convolutional transformer (TCT) is introduced for capturing complex temporal dynamics. To further optimize the model’s performance, we propose a novel temporal self-reference regularization (TSRR) loss to enhance temporal consistency and mitigate emotion ambiguity. Extensive evaluations demonstrate ATSLA-DFER achieves state-of-the-art (SOTA) performance on DFEW (69.48% UAR and 79.62% WAR) and FERV39k (44.89% UAR and 55.87% WAR), while achieving competitive performance on the challenging MAFW dataset (43.10% UAR and 56.24% WAR). ATSLA-DFER’s ability to effectively learn and integrate temporal–spatial features through its attention-driven architecture represents a significant step forward in advancing DFER capabilities for a wide range of practical applications in unconstrained environments.