Developing a Kazakh Audio–Visual Multimodal Speech Recognition Model Based on Hierarchical and Cross-Modal Attention
Information (Switzerland), cilt.17, sa.8, 2026 (ESCI, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 17 Sayı: 8
- Basım Tarihi: 2026
- Doi Numarası: 10.3390/info17080756
- Dergi Adı: Information (Switzerland)
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, Aerospace Database, Compendex, INSPEC, Library, Information Science & Technology Abstracts (LISTA), Directory of Open Access Journals, Information Science & Technology Abstracts (LISTA), Academic Search Ultimate (EBSCO), Technology Collection (ProQuest)
- Anahtar Kelimeler: AVSR, cross-modal attention, hierarchical attention, HuBERT, Kazakh language, multimodal speech recognition, Vision Transformer
- Gazi Üniversitesi Adresli: Evet
Özet
This study presents an audio–visual speech recognition (AVSR) model for Kazakh that jointly exploits audio and visual channels. The study introduces QazAVSR, a 57 h dataset collected from 271 speakers, and extracts synchronized audio signals and lip-region video sequences using FFmpeg 7.0, Dlib 19.24, and OpenCV 4.9.0. The proposed architecture uses the self-supervised HuBERT_BASE model in the audio branch and an ImageNet-pretrained ViT-B/16 model in the visual branch. Audio and visual representations are fused by a three-layer BiModalHformer block, where intra- and cross-attention operations are performed at each level. Extensive experimental validation, supplemented by rigorous paired bootstrap resampling significance tests, demonstrates that the full multimodal BiModalHformer model achieves a highly robust average character error rate (CER) of 31.2% and a Word Error Rate (WER) of 43.1%. These results significantly outperform traditional audio-only, video-only, and standard representation-level fusion baselines. Furthermore, comparisons against powerful external baseline architectures—including Whisper-Small and AV-HuBERT configurations rigorously adapted for the Kazakh language—statistically validate the architectural efficacy of the BiModalHformer framework. Additional systematic evaluations utilizing extended metrics such as the Match Error Rate (MER), word information preserved (WIP), and the Multimodal Synergy Index (MSI) confirm that the full audio–visual configuration preserves lexical information significantly more effectively. Finally, extensive noise perturbation experiments confirm that the multimodal architecture exhibits superior structural robustness to complex acoustic distortions, including environmental noise, synthetic room reverberation, and overlapping speech topologies.