OCR-Enhanced Multimodal ASR Can Read While Listening
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910000979378176 |
|---|---|
| author | Chen, Junli Tang, Changli Li, Yixuan Sun, Guangzhi Zhang, Chao |
| author_facet | Chen, Junli Tang, Changli Li, Yixuan Sun, Guangzhi Zhang, Chao |
| contents | Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18393 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | OCR-Enhanced Multimodal ASR Can Read While Listening Chen, Junli Tang, Changli Li, Yixuan Sun, Guangzhi Zhang, Chao Sound Computation and Language Audio and Speech Processing Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline. |
| title | OCR-Enhanced Multimodal ASR Can Read While Listening |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2601.18393 |