From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Haoyang, Hu, Yuchen, Chen, Chen, Siniscalchi, Sabato Marco, Liu, Songting, Chng, Eng Siong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912384645332992
author Li, Haoyang
Hu, Yuchen
Chen, Chen
Siniscalchi, Sabato Marco
Liu, Songting
Chng, Eng Siong
author_facet Li, Haoyang
Hu, Yuchen
Chen, Chen
Siniscalchi, Sabato Marco
Liu, Songting
Chng, Eng Siong
contents Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND show that GR-KAN requires up to 4x fewer parameters while improving PESQ by up to 0.1. In contrast, KAN, facing scalability issues, outperforms MLP on a small-scale signal modeling task but fails to improve MP-SENet. We demonstrate the first successful use of KAN-based methods for consistent improvement in both time- and SoTA TF-domain SE, establishing GR-KAN as a promising alternative for SE.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17778
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
Li, Haoyang
Hu, Yuchen
Chen, Chen
Siniscalchi, Sabato Marco
Liu, Songting
Chng, Eng Siong
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Deep neural network (DNN)-based speech enhancement (SE) usually uses conventional activation functions, which lack the expressiveness to capture complex multiscale structures needed for high-fidelity SE. Group-Rational KAN (GR-KAN), a variant of Kolmogorov-Arnold Networks (KAN), retains KAN's expressiveness while improving scalability on complex tasks. We adapt GR-KAN to existing DNN-based SE by replacing dense layers with GR-KAN layers in the time-frequency (T-F) domain MP-SENet and adapting GR-KAN's activations into the 1D CNN layers in the time-domain Demucs. Results on Voicebank-DEMAND show that GR-KAN requires up to 4x fewer parameters while improving PESQ by up to 0.1. In contrast, KAN, facing scalability issues, outperforms MLP on a small-scale signal modeling task but fails to improve MP-SENet. We demonstrate the first successful use of KAN-based methods for consistent improvement in both time- and SoTA TF-domain SE, establishing GR-KAN as a promising alternative for SE.
title From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.17778