OCR-Enhanced Multimodal ASR Can Read While Listening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junli, Tang, Changli, Li, Yixuan, Sun, Guangzhi, Zhang, Chao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910000979378176
author Chen, Junli
Tang, Changli
Li, Yixuan
Sun, Guangzhi
Zhang, Chao
author_facet Chen, Junli
Tang, Changli
Li, Yixuan
Sun, Guangzhi
Zhang, Chao
contents Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18393
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OCR-Enhanced Multimodal ASR Can Read While Listening
Chen, Junli
Tang, Changli
Li, Yixuan
Sun, Guangzhi
Zhang, Chao
Sound
Computation and Language
Audio and Speech Processing
Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline.
title OCR-Enhanced Multimodal ASR Can Read While Listening
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2601.18393