Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dehaghania, Parinaz Binandeh, Penab, Danilo, Aguiar, A. Pedro
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908849530732544
author Dehaghania, Parinaz Binandeh
Penab, Danilo
Aguiar, A. Pedro
author_facet Dehaghania, Parinaz Binandeh
Penab, Danilo
Aguiar, A. Pedro
contents Environmental sound classification (ESC) has gained significant attention due to its diverse applications in smart city monitoring, fault detection, acoustic surveillance, and manufacturing quality control. To enhance CNN performance, feature stacking techniques have been explored to aggregate complementary acoustic descriptors into richer input representations. In this paper, we investigate CNN-based models employing various stacked feature combinations, including Log-Mel Spectrogram (LM), Spectral Contrast (SPC), Chroma (CH), Tonnetz (TZ), Mel-Frequency Cepstral Coefficients (MFCCs), and Gammatone Cepstral Coefficients (GTCC). Experiments are conducted on the widely used ESC-50 and UrbanSound8K datasets under different training regimes, including pretraining on ESC-50, fine-tuning on UrbanSound8K, and comparison with Audio Spectrogram Transformer (AST) models pretrained on large-scale corpora such as AudioSet. This experimental design enables an analysis of how feature-stacked CNNs compare with transformer-based models under varying levels of training data and pretraining diversity. The results indicate that feature-stacked CNNs offer a more computationally and data-efficient alternative when large-scale pretraining or extensive training data are unavailable, making them particularly well suited for resource-constrained and edge-level sound classification scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
Dehaghania, Parinaz Binandeh
Penab, Danilo
Aguiar, A. Pedro
Audio and Speech Processing
Sound
Environmental sound classification (ESC) has gained significant attention due to its diverse applications in smart city monitoring, fault detection, acoustic surveillance, and manufacturing quality control. To enhance CNN performance, feature stacking techniques have been explored to aggregate complementary acoustic descriptors into richer input representations. In this paper, we investigate CNN-based models employing various stacked feature combinations, including Log-Mel Spectrogram (LM), Spectral Contrast (SPC), Chroma (CH), Tonnetz (TZ), Mel-Frequency Cepstral Coefficients (MFCCs), and Gammatone Cepstral Coefficients (GTCC). Experiments are conducted on the widely used ESC-50 and UrbanSound8K datasets under different training regimes, including pretraining on ESC-50, fine-tuning on UrbanSound8K, and comparison with Audio Spectrogram Transformer (AST) models pretrained on large-scale corpora such as AudioSet. This experimental design enables an analysis of how feature-stacked CNNs compare with transformer-based models under varying levels of training data and pretraining diversity. The results indicate that feature-stacked CNNs offer a more computationally and data-efficient alternative when large-scale pretraining or extensive training data are unavailable, making them particularly well suited for resource-constrained and edge-level sound classification scenarios.
title Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2602.09321