Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kuznetsova, Anastasia, Jang, Inseon, Lim, Wootaek, Kim, Minje
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916881783324672
author Kuznetsova, Anastasia
Jang, Inseon
Lim, Wootaek
Kim, Minje
author_facet Kuznetsova, Anastasia
Jang, Inseon
Lim, Wootaek
Kim, Minje
contents Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prioritizes efficient compression and downstream task performance, disregarding perceptual nuances. This work introduces an efficient ACoM method that can compress and quantize any chosen intermediate feature representation of an already trained speech/audio downstream model. Our approach employs task-specific loss guidance alongside residual vector quantization (RVQ) losses, providing ultra-low bitrates (i.e., less than 200 bps) with a minimal loss of the downstream model performance. The resulting tokenizer is adaptable to various bitrates and model sizes for flexible deployment. Evaluated on automatic speech recognition and audio classification, our method demonstrates its efficacy and potential for broader task and architectural applicability through appropriate regularization.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
Kuznetsova, Anastasia
Jang, Inseon
Lim, Wootaek
Kim, Minje
Sound
Artificial Intelligence
Audio and Speech Processing
Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prioritizes efficient compression and downstream task performance, disregarding perceptual nuances. This work introduces an efficient ACoM method that can compress and quantize any chosen intermediate feature representation of an already trained speech/audio downstream model. Our approach employs task-specific loss guidance alongside residual vector quantization (RVQ) losses, providing ultra-low bitrates (i.e., less than 200 bps) with a minimal loss of the downstream model performance. The resulting tokenizer is adaptable to various bitrates and model sizes for flexible deployment. Evaluated on automatic speech recognition and audio classification, our method demonstrates its efficacy and potential for broader task and architectural applicability through appropriate regularization.
title Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2507.12701