FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahasan, Md Mubtasim, Khan, Rafat Hasan, Mohiuddin, Tasnim, Chadha, Aman, Iqbal, Tariq, Amin, M Ashraful, Ali, Amin Ahsan, Islam, Md Mofijul, Rahman, A K M Mahbubur
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814368501760
author Ahasan, Md Mubtasim
Khan, Rafat Hasan
Mohiuddin, Tasnim
Chadha, Aman
Iqbal, Tariq
Amin, M Ashraful
Ali, Amin Ahsan
Islam, Md Mofijul
Rahman, A K M Mahbubur
author_facet Ahasan, Md Mubtasim
Khan, Rafat Hasan
Mohiuddin, Tasnim
Chadha, Aman
Iqbal, Tariq
Amin, M Ashraful
Ali, Amin Ahsan
Islam, Md Mofijul
Rahman, A K M Mahbubur
contents Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to human speech. While recent efforts introduced semantic representations from self-supervised speech models or incorporated contextual representations from pre-trained language models, challenges remain in aligning and unifying the semantic and contextual representations. We introduce FuseCodec, which unifies acoustic, semantic, and contextual representations through strong cross-modal alignment and globally informed supervision. We propose three complementary techniques: (i) Latent Representation Fusion, integrating semantic and contextual features directly into the encoder latent space for robust and unified representation learning; (ii) Global Semantic-Contextual Supervision, supervising discrete tokens with globally pooled and broadcasted representations to enhance temporal consistency and cross-modal alignment; and (iii) Temporally Aligned Contextual Supervision, strengthening alignment by dynamically matching contextual and speech tokens within a local window for fine-grained token-level supervision. We further introduce FuseCodec-TTS, demonstrating our methodology's applicability to zero-shot speech synthesis. Empirically, FuseCodec achieves state-of-the-art performance in LibriSpeech, surpassing EnCodec, SpeechTokenizer, and DAC in transcription accuracy, perceptual quality, intelligibility, and speaker similarity. Results highlight the effectiveness of contextually and semantically guided tokenization for speech tokenization and downstream tasks. Code and pretrained models are available at https://github.com/mubtasimahasan/FuseCodec.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11425
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
Ahasan, Md Mubtasim
Khan, Rafat Hasan
Mohiuddin, Tasnim
Chadha, Aman
Iqbal, Tariq
Amin, M Ashraful
Ali, Amin Ahsan
Islam, Md Mofijul
Rahman, A K M Mahbubur
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to human speech. While recent efforts introduced semantic representations from self-supervised speech models or incorporated contextual representations from pre-trained language models, challenges remain in aligning and unifying the semantic and contextual representations. We introduce FuseCodec, which unifies acoustic, semantic, and contextual representations through strong cross-modal alignment and globally informed supervision. We propose three complementary techniques: (i) Latent Representation Fusion, integrating semantic and contextual features directly into the encoder latent space for robust and unified representation learning; (ii) Global Semantic-Contextual Supervision, supervising discrete tokens with globally pooled and broadcasted representations to enhance temporal consistency and cross-modal alignment; and (iii) Temporally Aligned Contextual Supervision, strengthening alignment by dynamically matching contextual and speech tokens within a local window for fine-grained token-level supervision. We further introduce FuseCodec-TTS, demonstrating our methodology's applicability to zero-shot speech synthesis. Empirically, FuseCodec achieves state-of-the-art performance in LibriSpeech, surpassing EnCodec, SpeechTokenizer, and DAC in transcription accuracy, perceptual quality, intelligibility, and speaker similarity. Results highlight the effectiveness of contextually and semantically guided tokenization for speech tokenization and downstream tasks. Code and pretrained models are available at https://github.com/mubtasimahasan/FuseCodec.
title FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.11425