Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Sungnyun, Jang, Kangwook, Cho, Sungwoo, Chung, Joon Son, Kim, Hoirin, Yun, Se-Young
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908594301042688
author Kim, Sungnyun
Jang, Kangwook
Cho, Sungwoo
Chung, Joon Son
Kim, Hoirin
Yun, Se-Young
author_facet Kim, Sungnyun
Jang, Kangwook
Cho, Sungwoo
Chung, Joon Son
Kim, Hoirin
Yun, Se-Young
contents This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directly in the language space. Our framework, DualHyp, empowers a large language model (LLM) to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and visual speech recognition (VSR) models. To maximize the effectiveness of DualHyp, we further introduce RelPrompt, a noise-aware guidance mechanism that provides modality-grounded prompts to the LLM. RelPrompt offers the temporal reliability of each modality stream, guiding the model to dynamically switch its focus between ASR and VSR hypotheses for an accurate correction. Under various corruption scenarios, our framework attains up to 57.7% error rate gain on the LRS2 benchmark over standard ASR baseline, contrary to single-stream GER approaches that achieve only 10% gain. To facilitate research within our DualHyp framework, we release the code and the dataset comprising ASR and VSR hypotheses at https://github.com/sungnyun/dualhyp.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13281
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
Kim, Sungnyun
Jang, Kangwook
Cho, Sungwoo
Chung, Joon Son
Kim, Hoirin
Yun, Se-Young
Audio and Speech Processing
Computation and Language
Machine Learning
This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directly in the language space. Our framework, DualHyp, empowers a large language model (LLM) to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and visual speech recognition (VSR) models. To maximize the effectiveness of DualHyp, we further introduce RelPrompt, a noise-aware guidance mechanism that provides modality-grounded prompts to the LLM. RelPrompt offers the temporal reliability of each modality stream, guiding the model to dynamically switch its focus between ASR and VSR hypotheses for an accurate correction. Under various corruption scenarios, our framework attains up to 57.7% error rate gain on the LRS2 benchmark over standard ASR baseline, contrary to single-stream GER approaches that achieve only 10% gain. To facilitate research within our DualHyp framework, we release the code and the dataset comprising ASR and VSR hypotheses at https://github.com/sungnyun/dualhyp.
title Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
topic Audio and Speech Processing
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.13281