Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Casanova, Edresson, Langman, Ryan, Neekhara, Paarth, Hussain, Shehzeen, Li, Jason, Ghosh, Subhankar, Jukić, Ante, Lee, Sang-gil
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916399965798400
author Casanova, Edresson
Langman, Ryan
Neekhara, Paarth
Hussain, Shehzeen
Li, Jason
Ghosh, Subhankar
Jukić, Ante
Lee, Sang-gil
author_facet Casanova, Edresson
Langman, Ryan
Neekhara, Paarth
Hussain, Shehzeen
Li, Jason
Ghosh, Subhankar
Jukić, Ante
Lee, Sang-gil
contents Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12117
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
Casanova, Edresson
Langman, Ryan
Neekhara, Paarth
Hussain, Shehzeen
Li, Jason
Ghosh, Subhankar
Jukić, Ante
Lee, Sang-gil
Audio and Speech Processing
Computation and Language
Sound
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
title Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2409.12117