Neural Spectral Band Generation for Audio Coding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913961750822912 |
|---|---|
| author | Choi, Woongjib Kim, Byeong Hyeon Lim, Hyungseob Jang, Inseon Kang, Hong-Goo |
| author_facet | Choi, Woongjib Kim, Byeong Hyeon Lim, Hyungseob Jang, Inseon Kang, Hong-Goo |
| contents | Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06732 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Neural Spectral Band Generation for Audio Coding Choi, Woongjib Kim, Byeong Hyeon Lim, Hyungseob Jang, Inseon Kang, Hong-Goo Audio and Speech Processing Artificial Intelligence Signal Processing Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to diverse acoustic signals. In this paper, we explore the efficacy of a deep neural network (DNN)-based generative approach for coding the high-frequency bands, which we call neural spectral band generation (n-SBG). Specifically, we propose a DNN-based encoder-decoder structure to extract and quantize the side information related to the high-frequency components and generate the components given both the side information and the decoded core-band signals. The whole coding pipeline is optimized with generative adversarial criteria to enable the generation of perceptually plausible sound. From experiments using AAC as the core codec, we show that the proposed method achieves a better perceptual quality than HE-AAC-v1 with much less side information. |
| title | Neural Spectral Band Generation for Audio Coding |
| topic | Audio and Speech Processing Artificial Intelligence Signal Processing |
| url | https://arxiv.org/abs/2506.06732 |