MaLa-ASR: Multimedia-Assisted LLM-Based ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Guanrou, Ma, Ziyang, Yu, Fan, Gao, Zhifu, Zhang, Shiliang, Chen, Xie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912112712876032
author Yang, Guanrou
Ma, Ziyang
Yu, Fan
Gao, Zhifu
Zhang, Shiliang
Chen, Xie
author_facet Yang, Guanrou
Ma, Ziyang
Yu, Fan
Gao, Zhifu
Zhang, Shiliang
Chen, Xie
contents As more and more information-rich data like video become available, utilizing multi-modal auxiliary information to enhance audio tasks has sparked widespread research interest. The recent surge in research on LLM-based audio models provides fresh perspectives for tackling audio tasks. Given that LLM can flexibly ingest multiple inputs, we propose MaLa-ASR, an LLM-based ASR model that can integrate textual keywords extracted from presentation slides to improve recognition of conference content. MaLa-ASR yields average WERs of 9.4% and 11.7% on the L95 and S95 subsets of the SlideSpeech corpus, representing a significant relative WER drop of 27.9% and 44.7% over the baseline model reported in SlideSpeech. MaLa-ASR underscores LLM's strong performance in speech tasks and the capability to integrate auxiliary information conveniently. By adding keywords to the input prompt, the biased word error rate (B-WER) reduces relatively by 46.0% and 44.2%, establishing a new SOTA on this dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MaLa-ASR: Multimedia-Assisted LLM-Based ASR
Yang, Guanrou
Ma, Ziyang
Yu, Fan
Gao, Zhifu
Zhang, Shiliang
Chen, Xie
Audio and Speech Processing
Artificial Intelligence
As more and more information-rich data like video become available, utilizing multi-modal auxiliary information to enhance audio tasks has sparked widespread research interest. The recent surge in research on LLM-based audio models provides fresh perspectives for tackling audio tasks. Given that LLM can flexibly ingest multiple inputs, we propose MaLa-ASR, an LLM-based ASR model that can integrate textual keywords extracted from presentation slides to improve recognition of conference content. MaLa-ASR yields average WERs of 9.4% and 11.7% on the L95 and S95 subsets of the SlideSpeech corpus, representing a significant relative WER drop of 27.9% and 44.7% over the baseline model reported in SlideSpeech. MaLa-ASR underscores LLM's strong performance in speech tasks and the capability to integrate auxiliary information conveniently. By adding keywords to the input prompt, the biased word error rate (B-WER) reduces relatively by 46.0% and 44.2%, establishing a new SOTA on this dataset.
title MaLa-ASR: Multimedia-Assisted LLM-Based ASR
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2406.05839