Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, W. Ronny, Allauzen, Cyril, Chen, Tongzhou, Gupta, Kilol, Hu, Ke, Qin, James, Zhang, Yu, Wang, Yongqiang, Chang, Shuo-Yiin, Sainath, Tara N.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911763111346176
author Huang, W. Ronny
Allauzen, Cyril
Chen, Tongzhou
Gupta, Kilol
Hu, Ke
Qin, James
Zhang, Yu
Wang, Yongqiang
Chang, Shuo-Yiin
Sainath, Tara N.
author_facet Huang, W. Ronny
Allauzen, Cyril
Chen, Tongzhou
Gupta, Kilol
Hu, Ke
Qin, James
Zhang, Yu
Wang, Yongqiang
Chang, Shuo-Yiin
Sainath, Tara N.
contents In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12789
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
Huang, W. Ronny
Allauzen, Cyril
Chen, Tongzhou
Gupta, Kilol
Hu, Ke
Qin, James
Zhang, Yu
Wang, Yongqiang
Chang, Shuo-Yiin
Sainath, Tara N.
Computation and Language
Sound
Audio and Speech Processing
In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the parallelization capabilities of accelerator hardware. Our approach combines the Universal Speech Model (USM) and the PaLM 2 language model in per-segment scoring mode, achieving an average relative WER improvement across all languages of 10.8% on FLEURS and 3.6% on YouTube captioning. Furthermore, our comprehensive ablation study analyzes key parameters such as LLM size, context length, vocabulary size, fusion methodology. For instance, we explore the impact of LLM size ranging from 128M to 340B parameters on ASR performance. This study provides valuable insights into the factors influencing the effectiveness of practical large-scale LM-fused speech recognition systems.
title Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.12789