Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Su, Fei, Li, Cancan, Liu, Juan, Ju, Wei, Suo, Hongbin, Li, Ming
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911483894431744
author Su, Fei
Li, Cancan
Liu, Juan
Ju, Wei
Suo, Hongbin
Li, Ming
author_facet Su, Fei
Li, Cancan
Liu, Juan
Ju, Wei
Suo, Hongbin
Li, Ming
contents Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition performance and shown effectiveness for AVSR. However, prior approaches project audio and visual features independently or apply shallow fusion, limiting cross-modal alignment and complementary exchange while increasing the LLM's computational load. To address this, we propose AVUR-LLM, an LLM-based Audio-Visual Speech Recognition via Sparse Modality Alignment and Visual Unit-Guided Refinement. Experiments on LRS3 demonstrate state-of-the-art results for AVSR. Under additive-noise conditions at 0 dB SNR, it achieves 37% relative improvement over the baseline system.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03811
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
Su, Fei
Li, Cancan
Liu, Juan
Ju, Wei
Suo, Hongbin
Li, Ming
Sound
Multimedia
Audio and Speech Processing
Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition performance and shown effectiveness for AVSR. However, prior approaches project audio and visual features independently or apply shallow fusion, limiting cross-modal alignment and complementary exchange while increasing the LLM's computational load. To address this, we propose AVUR-LLM, an LLM-based Audio-Visual Speech Recognition via Sparse Modality Alignment and Visual Unit-Guided Refinement. Experiments on LRS3 demonstrate state-of-the-art results for AVSR. Under additive-noise conditions at 0 dB SNR, it achieves 37% relative improvement over the baseline system.
title Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2603.03811