FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Della Libera, Luca, Paissan, Francesco, Subakan, Cem, Ravanelli, Mirco
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912669325328384
author Della Libera, Luca
Paissan, Francesco
Subakan, Cem
Ravanelli, Mirco
author_facet Della Libera, Luca
Paissan, Francesco
Subakan, Cem
Ravanelli, Mirco
contents Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous audio into tokens using neural audio codecs. However, existing approaches face limitations, including high bitrates, the loss of either semantic or acoustic information, and the reliance on multi-codebook designs when trying to capture both, which increases architectural complexity for downstream tasks. To address these challenges, we introduce FocalCodec, an efficient low-bitrate codec based on focal modulation that utilizes a single binary codebook to compress speech between 0.16 and 0.65 kbps. FocalCodec delivers competitive performance in speech resynthesis and voice conversion at lower bitrates than the current state-of-the-art, while effectively handling multilingual speech and noisy environments. Evaluation on downstream tasks shows that FocalCodec successfully preserves sufficient semantic and acoustic information, while also being well-suited for generative modeling. Demo samples and code are available at https://lucadellalib.github.io/focalcodec-web/.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04465
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
Della Libera, Luca
Paissan, Francesco
Subakan, Cem
Ravanelli, Mirco
Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discretizing continuous audio into tokens using neural audio codecs. However, existing approaches face limitations, including high bitrates, the loss of either semantic or acoustic information, and the reliance on multi-codebook designs when trying to capture both, which increases architectural complexity for downstream tasks. To address these challenges, we introduce FocalCodec, an efficient low-bitrate codec based on focal modulation that utilizes a single binary codebook to compress speech between 0.16 and 0.65 kbps. FocalCodec delivers competitive performance in speech resynthesis and voice conversion at lower bitrates than the current state-of-the-art, while effectively handling multilingual speech and noisy environments. Evaluation on downstream tasks shows that FocalCodec successfully preserves sufficient semantic and acoustic information, while also being well-suited for generative modeling. Demo samples and code are available at https://lucadellalib.github.io/focalcodec-web/.
title FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
topic Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2502.04465