Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Phan, Dang Thoai, Huynh, Tuan Anh, Pham, Van Tuan, Tran, Cao Minh, Mai, Van Thuan, Tran, Ngoc Quy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909929695084544
author Phan, Dang Thoai
Huynh, Tuan Anh
Pham, Van Tuan
Tran, Cao Minh
Mai, Van Thuan
Tran, Ngoc Quy
author_facet Phan, Dang Thoai
Huynh, Tuan Anh
Pham, Van Tuan
Tran, Cao Minh
Mai, Van Thuan
Tran, Ngoc Quy
contents The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost poses a significant challenge, often leading researchers to prefer alternative methods such as the Short-Time Fourier Transform (STFT). To address this issue, this paper proposes a method to reduce the computational complexity of CWT by optimizing the length of the wavelet kernel and the hop size of the output scalogram. Experimental results demonstrate that the proposed approach significantly reduces computational cost while maintaining the robust performance of the trained model in acoustic recognition tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13017
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
Phan, Dang Thoai
Huynh, Tuan Anh
Pham, Van Tuan
Tran, Cao Minh
Mai, Van Thuan
Tran, Ngoc Quy
Audio and Speech Processing
Sound
Signal Processing
The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost poses a significant challenge, often leading researchers to prefer alternative methods such as the Short-Time Fourier Transform (STFT). To address this issue, this paper proposes a method to reduce the computational complexity of CWT by optimizing the length of the wavelet kernel and the hop size of the output scalogram. Experimental results demonstrate that the proposed approach significantly reduces computational cost while maintaining the robust performance of the trained model in acoustic recognition tasks.
title Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2505.13017