A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion
Athira Raj, Christy James Jose, K. S. BijuSpeech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces a multimodal SER model for Malayalam that uses both speech and text information. Because few Malayalam emotion databases are available, a dataset was constructed from publicly available audiovisual content. The method uses a wav2vec 2.0 model to extract audio features, while the text features are obtained with a pretrained language model. The unified model then fuses the audio and text features for emotion prediction. Label consistency in the constructed dataset was further examined with an embedding-based analysis. On a dataset of 1,040 samples evenly distributed across four emotions, the proposed model achieved an accuracy of 78.85% and a macro-averaged F1-score of 0.79 on the test set, outperforming an audio-only baseline.