Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prabhavalkar, Rohit, Meng, Zhong, Wang, Weiran, Stooke, Adam, Cai, Xingyu, He, Yanzhang, Narayanan, Arun, Hwang, Dongseong, Sainath, Tara N., Moreno, Pedro J.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917599011405824
author Prabhavalkar, Rohit
Meng, Zhong
Wang, Weiran
Stooke, Adam
Cai, Xingyu
He, Yanzhang
Narayanan, Arun
Hwang, Dongseong
Sainath, Tara N.
Moreno, Pedro J.
author_facet Prabhavalkar, Rohit
Meng, Zhong
Wang, Weiran
Stooke, Adam
Cai, Xingyu
He, Yanzhang
Narayanan, Arun
Hwang, Dongseong
Sainath, Tara N.
Moreno, Pedro J.
contents The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2402_17184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
Prabhavalkar, Rohit
Meng, Zhong
Wang, Weiran
Stooke, Adam
Cai, Xingyu
He, Yanzhang
Narayanan, Arun
Hwang, Dongseong
Sainath, Tara N.
Moreno, Pedro J.
Computation and Language
Sound
Audio and Speech Processing
The accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline.
title Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2402.17184