Bringing Interpretability to Neural Audio Codecs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadok, Samir, Hauret, Julien, Bavu, Éric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912592354607104
author Sadok, Samir
Hauret, Julien
Bavu, Éric
author_facet Sadok, Samir
Hauret, Julien
Bavu, Éric
contents The advent of neural audio codecs has increased in popularity due to their potential for efficiently modeling audio with transformers. Such advanced codecs represent audio from a highly continuous waveform to low-sampled discrete units. In contrast to semantic units, acoustic units may lack interpretability because their training objectives primarily focus on reconstruction performance. This paper proposes a two-step approach to explore the encoding of speech information within the codec tokens. The primary goal of the analysis stage is to gain deeper insight into how speech attributes such as content, identity, and pitch are encoded. The synthesis stage then trains an AnCoGen network for post-hoc explanation of codecs to extract speech attributes from the respective tokens directly.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bringing Interpretability to Neural Audio Codecs
Sadok, Samir
Hauret, Julien
Bavu, Éric
Audio and Speech Processing
The advent of neural audio codecs has increased in popularity due to their potential for efficiently modeling audio with transformers. Such advanced codecs represent audio from a highly continuous waveform to low-sampled discrete units. In contrast to semantic units, acoustic units may lack interpretability because their training objectives primarily focus on reconstruction performance. This paper proposes a two-step approach to explore the encoding of speech information within the codec tokens. The primary goal of the analysis stage is to gain deeper insight into how speech attributes such as content, identity, and pitch are encoded. The synthesis stage then trains an AnCoGen network for post-hoc explanation of codecs to extract speech attributes from the respective tokens directly.
title Bringing Interpretability to Neural Audio Codecs
topic Audio and Speech Processing
url https://arxiv.org/abs/2506.04492