Neural Spectral Band Generation for Audio Coding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Choi, Woongjib, Kim, Byeong Hyeon, Lim, Hyungseob, Jang, Inseon, Kang, Hong-Goo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913961750822912
author Choi, Woongjib
Kim, Byeong Hyeon
Lim, Hyungseob
Jang, Inseon
Kang, Hong-Goo
author_facet Choi, Woongjib
Kim, Byeong Hyeon
Lim, Hyungseob
Jang, Inseon
Kang, Hong-Goo
contents Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neural Spectral Band Generation for Audio Coding
Choi, Woongjib
Kim, Byeong Hyeon
Lim, Hyungseob
Jang, Inseon
Kang, Hong-Goo
Audio and Speech Processing
Artificial Intelligence
Signal Processing
Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information.
title Neural Spectral Band Generation for Audio Coding
topic Audio and Speech Processing
Artificial Intelligence
Signal Processing
url https://arxiv.org/abs/2506.06732