Learner emotion classification in online English instruction based on multimodal analysis
Xiaoli Huang, Adeel Ashraf CheemaThe virtualized nature of online English instruction weakens affective interaction between teachers and learners. This study proposes a multimodal dialogue emotion recognition method based on a semantic-enhanced graph and develops an emotion classification framework optimized for the unique multimodal noise and implicit learner emotional expression characteristics of online English teaching environments. Textual, acoustic, and visual features are first extracted using Robustly Optimized BERT Pretraining Approach (RoBERTa), OpenSmile, and DenseNet, respectively, and temporal and contextual dependencies are preserved using a bidirectional gated recurrent unit (BiGRU). The stochastic uncertainty of unimodal features is then quantified, and contrastive self-supervised learning is introduced to improve feature consistency and reduce the influence of environmental noise. Subsequently, a semantic-enhanced graph is constructed according to speaker dependency relationships. By integrating graph topology with a multi-head attention mechanism, a semantic information refinement module is designed to capture local and global emotional dependencies within dialogues jointly. Experiments conducted on the CMU-MOSEI and MS COCO datasets show that the proposed model achieves 78.9% accuracy and an F1-score (F1) of 0.76. Compared with the best-performing unimodal model, the accuracy and F1 improve by 8.8 and 0.09, respectively, and by 3.6 and 0.04 compared with an early-fusion BiGRU model. For implicit cross-modal emotional expressions, the recall rate improves by 18.7% relative to the baseline model. In terms of computational efficiency, the training times for the two datasets are 4.1 and 4.3 h, respectively.