FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Della Libera, Luca, Subakan, Cem, Ravanelli, Mirco
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916958387044352
author Della Libera, Luca
Subakan, Cem
Ravanelli, Mirco
author_facet Della Libera, Luca
Subakan, Cem
Ravanelli, Mirco
contents Neural audio codecs are a fundamental component of modern generative audio pipelines. Although recent codecs achieve strong low-bitrate reconstruction and provide powerful representations for downstream tasks, most are non-streamable, limiting their use in real-time applications. We present FocalCodec-Stream, a hybrid codec based on focal modulation that compresses speech into a single binary codebook at 0.55 - 0.80 kbps with a theoretical latency of 80 ms. Our approach combines multi-stage causal distillation of WavLM with targeted architectural improvements, including a lightweight refiner module that enhances quality under latency constraints. Experiments show that FocalCodec-Stream outperforms existing streamable codecs at comparable bitrates, while preserving both semantic and acoustic information. The result is a favorable trade-off between reconstruction quality, downstream task performance, latency, and efficiency. Code and checkpoints will be released at https://github.com/lucadellalib/focalcodec.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
Della Libera, Luca
Subakan, Cem
Ravanelli, Mirco
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Neural audio codecs are a fundamental component of modern generative audio pipelines. Although recent codecs achieve strong low-bitrate reconstruction and provide powerful representations for downstream tasks, most are non-streamable, limiting their use in real-time applications. We present FocalCodec-Stream, a hybrid codec based on focal modulation that compresses speech into a single binary codebook at 0.55 - 0.80 kbps with a theoretical latency of 80 ms. Our approach combines multi-stage causal distillation of WavLM with targeted architectural improvements, including a lightweight refiner module that enhances quality under latency constraints. Experiments show that FocalCodec-Stream outperforms existing streamable codecs at comparable bitrates, while preserving both semantic and acoustic information. The result is a favorable trade-off between reconstruction quality, downstream task performance, latency, and efficiency. Code and checkpoints will be released at https://github.com/lucadellalib/focalcodec.
title FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.16195