DOI: 10.3390/info17080756 ISSN: 2078-2489

Developing a Kazakh Audio–Visual Multimodal Speech Recognition Model Based on Hierarchical and Cross-Modal Attention

Turdybek Kurmetkan, Orken Mamyrbayev, Adem Tekerek, Ainur Toleu

This study presents an audio–visual speech recognition (AVSR) model for Kazakh that jointly exploits audio and visual channels. The study introduces QazAVSR, a 57 h dataset collected from 271 speakers, and extracts synchronized audio signals and lip-region video sequences using FFmpeg 7.0, Dlib 19.24, and OpenCV 4.9.0. The proposed architecture uses the self-supervised HuBERT_BASE model in the audio branch and an ImageNet-pretrained ViT-B/16 model in the visual branch. Audio and visual representations are fused by a three-layer BiModalHformer block, where intra- and cross-attention operations are performed at each level. Extensive experimental validation, supplemented by rigorous paired bootstrap resampling significance tests, demonstrates that the full multimodal BiModalHformer model achieves a highly robust average character error rate (CER) of 31.2% and a Word Error Rate (WER) of 43.1%. These results significantly outperform traditional audio-only, video-only, and standard representation-level fusion baselines. Furthermore, comparisons against powerful external baseline architectures—including Whisper-Small and AV-HuBERT configurations rigorously adapted for the Kazakh language—statistically validate the architectural efficacy of the BiModalHformer framework. Additional systematic evaluations utilizing extended metrics such as the Match Error Rate (MER), word information preserved (WIP), and the Multimodal Synergy Index (MSI) confirm that the full audio–visual configuration preserves lexical information significantly more effectively. Finally, extensive noise perturbation experiments confirm that the multimodal architecture exhibits superior structural robustness to complex acoustic distortions, including environmental noise, synthetic room reverberation, and overlapping speech topologies.

More from our Archive