DOI: 10.68337/cpsm.v1.i1.2026-012 ISSN:

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Athira Raj, Christy James Jose, K. S. Biju

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces a multimodal SER model for Malayalam that uses both speech and text information. Because few Malayalam emotion databases are available, a dataset was constructed from publicly available audiovisual content. The method uses a wav2vec 2.0 model to extract audio features, while the text features are obtained with a pretrained language model. The unified model then fuses the audio and text features for emotion prediction. Label consistency in the constructed dataset was further examined with an embedding-based analysis. On a dataset of 1,040 samples evenly distributed across four emotions, the proposed model achieved an accuracy of 78.85% and a macro-averaged F1-score of 0.79 on the test set, outperforming an audio-only baseline.