VHASR: A Multimodal Speech Recognition System With Vision Hotwords

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jiliang, Li, Zuchao, Wang, Ping, Ai, Haojun, Zhang, Lefei, Zhao, Hai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910634223861760
author Hu, Jiliang
Li, Zuchao
Wang, Ping
Ai, Haojun
Zhang, Lefei
Zhao, Hai
author_facet Hu, Jiliang
Li, Zuchao
Wang, Ping
Ai, Haojun
Zhang, Lefei
Zhao, Hai
contents The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that introducing image information to model does not help improving ASR performance. In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability. Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs. We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages. The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability. Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00822
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VHASR: A Multimodal Speech Recognition System With Vision Hotwords
Hu, Jiliang
Li, Zuchao
Wang, Ping
Ai, Haojun
Zhang, Lefei
Zhao, Hai
Sound
Computation and Language
Audio and Speech Processing
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that introducing image information to model does not help improving ASR performance. In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability. Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs. We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages. The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability. Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR.
title VHASR: A Multimodal Speech Recognition System With Vision Hotwords
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2410.00822