Trainingless Adaptation of Pretrained Models for Environmental Sound Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tonami, Noriyuki, Kohno, Wataru, Imoto, Keisuke, Yajima, Yoshiyuki, Mishima, Sakiko, Kondo, Reishi, Hino, Tomoyuki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915075883794432
author Tonami, Noriyuki
Kohno, Wataru
Imoto, Keisuke
Yajima, Yoshiyuki
Mishima, Sakiko
Kondo, Reishi
Hino, Tomoyuki
author_facet Tonami, Noriyuki
Kohno, Wataru
Imoto, Keisuke
Yajima, Yoshiyuki
Mishima, Sakiko
Kondo, Reishi
Hino, Tomoyuki
contents Deep neural network (DNN)-based models for environmental sound classification are not robust against a domain to which training data do not belong, that is, out-of-distribution or unseen data. To utilize pretrained models for the unseen domain, adaptation methods, such as finetuning and transfer learning, are used with rich computing resources, e.g., the graphical processing unit (GPU). However, it is becoming more difficult to keep up with research trends for those who have poor computing resources because state-of-the-art models are becoming computationally resource-intensive. In this paper, we propose a trainingless adaptation method for pretrained models for environmental sound classification. To introduce the trainingless adaptation method, we first propose an operation of recovering time--frequency-ish (TF-ish) structures in intermediate layers of DNN models. We then propose the trainingless frequency filtering method for domain adaptation, which is not a gradient-based optimization widely used. The experiments conducted using the ESC-50 dataset show that the proposed adaptation method improves the classification accuracy by 20.40 percentage points compared with the conventional method.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17212
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
Tonami, Noriyuki
Kohno, Wataru
Imoto, Keisuke
Yajima, Yoshiyuki
Mishima, Sakiko
Kondo, Reishi
Hino, Tomoyuki
Sound
Audio and Speech Processing
Deep neural network (DNN)-based models for environmental sound classification are not robust against a domain to which training data do not belong, that is, out-of-distribution or unseen data. To utilize pretrained models for the unseen domain, adaptation methods, such as finetuning and transfer learning, are used with rich computing resources, e.g., the graphical processing unit (GPU). However, it is becoming more difficult to keep up with research trends for those who have poor computing resources because state-of-the-art models are becoming computationally resource-intensive. In this paper, we propose a trainingless adaptation method for pretrained models for environmental sound classification. To introduce the trainingless adaptation method, we first propose an operation of recovering time--frequency-ish (TF-ish) structures in intermediate layers of DNN models. We then propose the trainingless frequency filtering method for domain adaptation, which is not a gradient-based optimization widely used. The experiments conducted using the ESC-50 dataset show that the proposed adaptation method improves the classification accuracy by 20.40 percentage points compared with the conventional method.
title Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.17212