A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Dingdong, Cui, Mingyu, Yang, Dongchao, Chen, Xueyuan, Meng, Helen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917836187762688
author Wang, Dingdong
Cui, Mingyu
Yang, Dongchao
Chen, Xueyuan
Meng, Helen
author_facet Wang, Dingdong
Cui, Mingyu
Yang, Dongchao
Chen, Xueyuan
Meng, Helen
contents With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamlessly. Compared to most studies that focus on continuous speech features, although discrete-token based LLMs have shown promising results on certain tasks, the performance gap between these two paradigms is rarely explored. In this paper, we present a fair and thorough comparison between discrete and continuous features across a variety of semantic-related tasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal that continuous features generally outperform discrete tokens, particularly in tasks requiring fine-grained semantic understanding. Moreover, this study goes beyond surface-level comparison by identifying key factors behind the under-performance of discrete tokens, such as limited token granularity and inefficient information retention. To enhance the performance of discrete tokens, we explore potential aspects based on our analysis. We hope our results can offer new insights into the opportunities for advancing discrete speech tokens in Speech LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08742
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
Wang, Dingdong
Cui, Mingyu
Yang, Dongchao
Chen, Xueyuan
Meng, Helen
Computation and Language
Sound
Audio and Speech Processing
With the rise of Speech Large Language Models (Speech LLMs), there has been growing interest in discrete speech tokens for their ability to integrate with text-based tokens seamlessly. Compared to most studies that focus on continuous speech features, although discrete-token based LLMs have shown promising results on certain tasks, the performance gap between these two paradigms is rarely explored. In this paper, we present a fair and thorough comparison between discrete and continuous features across a variety of semantic-related tasks using a light-weight LLM (Qwen1.5-0.5B). Our findings reveal that continuous features generally outperform discrete tokens, particularly in tasks requiring fine-grained semantic understanding. Moreover, this study goes beyond surface-level comparison by identifying key factors behind the under-performance of discrete tokens, such as limited token granularity and inefficient information retention. To enhance the performance of discrete tokens, we explore potential aspects based on our analysis. We hope our results can offer new insights into the opportunities for advancing discrete speech tokens in Speech LLMs.
title A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.08742