Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shechtman, Slava, Dekel, Avihu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914969598033920
author Shechtman, Slava
Dekel, Avihu
author_facet Shechtman, Slava
Dekel, Avihu
contents Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. Various publicly available trainable discrete tokenizers recently demonstrated impressive results for audio tokenization, yet they mostly require high token rates to gain high-quality reconstruction. In this study, we fine-tuned an open-source general audio RVQGAN model using diverse open-source speech data, considering various recording conditions and quality levels. The resulting wideband (24kHz) speech-only model achieves speech reconstruction, which is nearly indistinguishable from PCM (pulse-code modulation) with a rate of 150-300 tokens per second (1500-3000 bps). The evaluation used comprehensive English speech data encompassing different recording conditions, including studio settings. Speech samples are made publicly available in http://ibm.biz/IS24SpeechRVQ . The model is officially released in https://huggingface.co/ibm/DAC.speech.v1.0
format Preprint
id arxiv_https___arxiv_org_abs_2410_08325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
Shechtman, Slava
Dekel, Avihu
Audio and Speech Processing
Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. Various publicly available trainable discrete tokenizers recently demonstrated impressive results for audio tokenization, yet they mostly require high token rates to gain high-quality reconstruction. In this study, we fine-tuned an open-source general audio RVQGAN model using diverse open-source speech data, considering various recording conditions and quality levels. The resulting wideband (24kHz) speech-only model achieves speech reconstruction, which is nearly indistinguishable from PCM (pulse-code modulation) with a rate of 150-300 tokens per second (1500-3000 bps). The evaluation used comprehensive English speech data encompassing different recording conditions, including studio settings. Speech samples are made publicly available in http://ibm.biz/IS24SpeechRVQ . The model is officially released in https://huggingface.co/ibm/DAC.speech.v1.0
title Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
topic Audio and Speech Processing
url https://arxiv.org/abs/2410.08325