As Good as It KAN Get: High-Fidelity Audio Representation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Marszałek, Patryk, Rut, Maciej, Kawa, Piotr, Spurek, Przemysław, Syga, Piotr
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914128768008192
author Marszałek, Patryk
Rut, Maciej
Kawa, Piotr
Spurek, Przemysław
Syga, Piotr
author_facet Marszałek, Patryk
Rut, Maciej
Kawa, Piotr
Spurek, Przemysław
Syga, Piotr
contents Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-SpectralDistance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks. The source code can be accessed at https://github.com/gmum/fewsound.git.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02585
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle As Good as It KAN Get: High-Fidelity Audio Representation
Marszałek, Patryk
Rut, Maciej
Kawa, Piotr
Spurek, Przemysław
Syga, Piotr
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-SpectralDistance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks. The source code can be accessed at https://github.com/gmum/fewsound.git.
title As Good as It KAN Get: High-Fidelity Audio Representation
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2503.02585