Factorized RVQ-GAN For Disentangled Speech Tokenization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Khurana, Sameer, Klement, Dominik, Laurent, Antoine, Bobos, Dominik, Novosad, Juraj, Gazdik, Peter, Zhang, Ellen, Huang, Zili, Hussein, Amir, Marxer, Ricard, Masuyama, Yoshiki, Aihara, Ryo, Hori, Chiori, Germain, Francois G., Wichern, Gordon, Roux, Jonathan Le
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912438198206464
author Khurana, Sameer
Klement, Dominik
Laurent, Antoine
Bobos, Dominik
Novosad, Juraj
Gazdik, Peter
Zhang, Ellen
Huang, Zili
Hussein, Amir
Marxer, Ricard
Masuyama, Yoshiki
Aihara, Ryo
Hori, Chiori
Germain, Francois G.
Wichern, Gordon
Roux, Jonathan Le
author_facet Khurana, Sameer
Klement, Dominik
Laurent, Antoine
Bobos, Dominik
Novosad, Juraj
Gazdik, Peter
Zhang, Ellen
Huang, Zili
Hussein, Amir
Marxer, Ricard
Masuyama, Yoshiki
Aihara, Ryo
Hori, Chiori
Germain, Francois G.
Wichern, Gordon
Roux, Jonathan Le
contents We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives: one from a pre-trained speech encoder (HuBERT) for phoneme-level structure, and another from a text-based encoder (LaBSE) for lexical cues. Experiments on English and multilingual data show that HAC's factorized bottleneck yields disentangled token sets: one aligns with phonemes, while another captures word-level semantics. Quantitative evaluations confirm that HAC tokens preserve naturalness and provide interpretable linguistic information, outperforming single-level baselines in both disentanglement and reconstruction quality. These findings underscore HAC's potential as a unified discrete speech representation, bridging acoustic detail and lexical meaning for downstream speech generation and understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Factorized RVQ-GAN For Disentangled Speech Tokenization
Khurana, Sameer
Klement, Dominik
Laurent, Antoine
Bobos, Dominik
Novosad, Juraj
Gazdik, Peter
Zhang, Ellen
Huang, Zili
Hussein, Amir
Marxer, Ricard
Masuyama, Yoshiki
Aihara, Ryo
Hori, Chiori
Germain, Francois G.
Wichern, Gordon
Roux, Jonathan Le
Audio and Speech Processing
Computation and Language
Sound
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives: one from a pre-trained speech encoder (HuBERT) for phoneme-level structure, and another from a text-based encoder (LaBSE) for lexical cues. Experiments on English and multilingual data show that HAC's factorized bottleneck yields disentangled token sets: one aligns with phonemes, while another captures word-level semantics. Quantitative evaluations confirm that HAC tokens preserve naturalness and provide interpretable linguistic information, outperforming single-level baselines in both disentanglement and reconstruction quality. These findings underscore HAC's potential as a unified discrete speech representation, bridging acoustic detail and lexical meaning for downstream speech generation and understanding tasks.
title Factorized RVQ-GAN For Disentangled Speech Tokenization
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2506.15456