Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qian, Wang, Wen, Zhang, Qinglin, Zheng, Siqi, Zhang, Shiliang, Deng, Chong, Ma, Yukun, Yu, Hai, Liu, Jiaqing, Zhang, Chong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916114094620672
author Chen, Qian
Wang, Wen
Zhang, Qinglin
Zheng, Siqi
Zhang, Shiliang
Deng, Chong
Ma, Yukun
Yu, Hai
Liu, Jiaqing
Zhang, Chong
author_facet Chen, Qian
Wang, Wen
Zhang, Qinglin
Zheng, Siqi
Zhang, Shiliang
Deng, Chong
Ma, Yukun
Yu, Hai
Liu, Jiaqing
Zhang, Chong
contents Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods. The source code can be found here: https://github.com/alibaba-damo-academy/SpokenNLP/tree/main/sld
format Preprint
id arxiv_https___arxiv_org_abs_2311_04534
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
Chen, Qian
Wang, Wen
Zhang, Qinglin
Zheng, Siqi
Zhang, Shiliang
Deng, Chong
Ma, Yukun
Yu, Hai
Liu, Jiaqing
Zhang, Chong
Computation and Language
Sound
Audio and Speech Processing
Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods. The source code can be found here: https://github.com/alibaba-damo-academy/SpokenNLP/tree/main/sld
title Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2311.04534