NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Casanova, Edresson, Neekhara, Paarth, Langman, Ryan, Hussain, Shehzeen, Ghosh, Subhankar, Yang, Xuesong, Jukić, Ante, Li, Jason, Ginsburg, Boris
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911097848594432
author Casanova, Edresson
Neekhara, Paarth
Langman, Ryan
Hussain, Shehzeen
Ghosh, Subhankar
Yang, Xuesong
Jukić, Ante
Li, Jason
Ginsburg, Boris
author_facet Casanova, Edresson
Neekhara, Paarth
Langman, Ryan
Hussain, Shehzeen
Ghosh, Subhankar
Yang, Xuesong
Jukić, Ante
Li, Jason
Ginsburg, Boris
contents Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling techniques to speech data. However, existing audio codecs often operate at high frame rates, leading to slow training and inference, particularly for autoregressive models. To address this, there is growing interest in low frame-rate audio codecs, which reduce the number of autoregressive steps required to generate one second of audio. In this paper, we conduct ablation studies to examine the impact of frame rate, bitrate, and causality on codec reconstruction quality. Based on our findings, we introduce NanoCodec, a state-of-the-art audio codec that achieves high-quality compression at just 12.5 frames per second (FPS). NanoCodec outperforms related works across various bitrate ranges, establishing a new benchmark for low-latency and efficient Speech LLM training and inference.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Casanova, Edresson
Neekhara, Paarth
Langman, Ryan
Hussain, Shehzeen
Ghosh, Subhankar
Yang, Xuesong
Jukić, Ante
Li, Jason
Ginsburg, Boris
Audio and Speech Processing
Computation and Language
Sound
Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling techniques to speech data. However, existing audio codecs often operate at high frame rates, leading to slow training and inference, particularly for autoregressive models. To address this, there is growing interest in low frame-rate audio codecs, which reduce the number of autoregressive steps required to generate one second of audio. In this paper, we conduct ablation studies to examine the impact of frame rate, bitrate, and causality on codec reconstruction quality. Based on our findings, we introduce NanoCodec, a state-of-the-art audio codec that achieves high-quality compression at just 12.5 frames per second (FPS). NanoCodec outperforms related works across various bitrate ranges, establishing a new benchmark for low-latency and efficient Speech LLM training and inference.
title NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2508.05835