Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lim, DongHoon, Kim, YoungChae, Kim, Dong-Hyun, Yang, Da-Hee, Chang, Joon-Hyuk
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908503943151616
author Lim, DongHoon
Kim, YoungChae
Kim, Dong-Hyun
Yang, Da-Hee
Chang, Joon-Hyuk
author_facet Lim, DongHoon
Kim, YoungChae
Kim, Dong-Hyun
Yang, Da-Hee
Chang, Joon-Hyuk
contents Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
Lim, DongHoon
Kim, YoungChae
Kim, Dong-Hyun
Yang, Da-Hee
Chang, Joon-Hyuk
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Audio and Speech Processing
Signal Processing
Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise.
title Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2508.18734