SUNAC: Source-aware Unified Neural Audio Codec

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aihara, Ryo, Masuyama, Yoshiki, Paissan, Francesco, Germain, François G., Wichern, Gordon, Roux, Jonathan Le
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908666681098240
author Aihara, Ryo
Masuyama, Yoshiki
Paissan, Francesco
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
author_facet Aihara, Ryo
Masuyama, Yoshiki
Paissan, Francesco
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
contents Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need access to only a subset of the sources (e.g., analysis of a particular type of sound, transcription of a given speaker, etc). To address this, we propose a source-aware codec that encodes individual sources directly from mixtures, conditioned on source type prompts. This enables user-driven selection of which source(s) to encode, including separately encoding multiple sources of the same type (e.g., multiple speech signals). Experiments show that our model achieves competitive resynthesis and separation quality relative to a cascade of source separation followed by a conventional NAC, with lower computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16126
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SUNAC: Source-aware Unified Neural Audio Codec
Aihara, Ryo
Masuyama, Yoshiki
Paissan, Francesco
Germain, François G.
Wichern, Gordon
Roux, Jonathan Le
Audio and Speech Processing
Signal Processing
Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need access to only a subset of the sources (e.g., analysis of a particular type of sound, transcription of a given speaker, etc). To address this, we propose a source-aware codec that encodes individual sources directly from mixtures, conditioned on source type prompts. This enables user-driven selection of which source(s) to encode, including separately encoding multiple sources of the same type (e.g., multiple speech signals). Experiments show that our model achieves competitive resynthesis and separation quality relative to a cascade of source separation followed by a conventional NAC, with lower computational cost.
title SUNAC: Source-aware Unified Neural Audio Codec
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2511.16126