LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kwak, Doyeop, Choi, Jeongsoo, Lee, Suyeon, Chung, Joon Son
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914520984715264
author Kwak, Doyeop
Choi, Jeongsoo
Lee, Suyeon
Chung, Joon Son
author_facet Kwak, Doyeop
Choi, Jeongsoo
Lee, Suyeon
Chung, Joon Son
contents We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select AVSR-suitable samples and preprocess them in an LRS-style format for direct use in existing AVSR pipelines. Compared with commonly used benchmarks, LRS-VoxMM covers a more diverse range of scenarios and acoustic conditions. We also release distorted evaluation sets with additive noise, reverberation, and bandwidth limitation to support evaluation under severe acoustic degradation. Experimental results show that LRS-VoxMM is considerably harder than LRS3 and that the contribution of visual information becomes more evident as the audio signal degrades. LRS-VoxMM supports more realistic AVSR benchmarking and encourages further research on the role of visual information in challenging real-world conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27866
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
Kwak, Doyeop
Choi, Jeongsoo
Lee, Suyeon
Chung, Joon Son
Audio and Speech Processing
Multimedia
Sound
We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select AVSR-suitable samples and preprocess them in an LRS-style format for direct use in existing AVSR pipelines. Compared with commonly used benchmarks, LRS-VoxMM covers a more diverse range of scenarios and acoustic conditions. We also release distorted evaluation sets with additive noise, reverberation, and bandwidth limitation to support evaluation under severe acoustic degradation. Experimental results show that LRS-VoxMM is considerably harder than LRS3 and that the contribution of visual information becomes more evident as the audio signal degrades. LRS-VoxMM supports more realistic AVSR benchmarking and encourages further research on the role of visual information in challenging real-world conditions.
title LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
topic Audio and Speech Processing
Multimedia
Sound
url https://arxiv.org/abs/2604.27866