Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imamura, Kanami, Nakamura, Tomohiko, Yatabe, Kohei, Saruwatari, Hiroshi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914268981493760
author Imamura, Kanami
Nakamura, Tomohiko
Yatabe, Kohei
Saruwatari, Hiroshi
author_facet Imamura, Kanami
Nakamura, Tomohiko
Yatabe, Kohei
Saruwatari, Hiroshi
contents Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the causes of this degradation through two hypotheses: (i) the lack of high-frequency components introduced by up-sampling, and (ii) the greater importance of their presence than their precise representation. To examine these hypotheses, we compare conventional resampling with three alternatives: post-resampling noise addition, which adds Gaussian noise to the resampled signal; noisy-kernel resampling, which perturbs the kernel with Gaussian noise to enrich high-frequency components; and trainable-kernel resampling, which adapts the interpolation kernel through training. Experiments on music source separation show that noisy-kernel and trainable-kernel resampling alleviate the degradation observed with conventional resampling. We further demonstrate that noisy-kernel resampling is effective across diverse models, highlighting it as a simple yet practical option.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14684
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch
Imamura, Kanami
Nakamura, Tomohiko
Yatabe, Kohei
Saruwatari, Hiroshi
Sound
Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the causes of this degradation through two hypotheses: (i) the lack of high-frequency components introduced by up-sampling, and (ii) the greater importance of their presence than their precise representation. To examine these hypotheses, we compare conventional resampling with three alternatives: post-resampling noise addition, which adds Gaussian noise to the resampled signal; noisy-kernel resampling, which perturbs the kernel with Gaussian noise to enrich high-frequency components; and trainable-kernel resampling, which adapts the interpolation kernel through training. Experiments on music source separation show that noisy-kernel and trainable-kernel resampling alleviate the degradation observed with conventional resampling. We further demonstrate that noisy-kernel resampling is effective across diverse models, highlighting it as a simple yet practical option.
title Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch
topic Sound
url https://arxiv.org/abs/2601.14684