Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ritu, Jarin, Mohammadi, Amirmohammad, Carreiro, Davelle, Van Dine, Alexandra, Peeples, Joshua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912979363037184
author Ritu, Jarin
Mohammadi, Amirmohammad
Carreiro, Davelle
Van Dine, Alexandra
Peeples, Joshua
author_facet Ritu, Jarin
Mohammadi, Amirmohammad
Carreiro, Davelle
Van Dine, Alexandra
Peeples, Joshua
contents While knowledge distillation has shown success in various audio tasks, its application to environmental sound classification often overlooks essential low-level audio texture features needed to capture local patterns in complex acoustic environments. To address this gap, the Structural and Statistical Audio Texture Knowledge Distillation (SSATKD) framework is proposed, which combines high-level contextual information with low-level structural and statistical audio textures extracted from intermediate layers. To evaluate its generalizability across diverse acoustic domains, SSATKD is tested on four datasets within the environmental sound classification domain, including two passive sonar datasets (DeepShip and Vessel Type Underwater Acoustic Data (VTUAD)) and two general environmental sound datasets (Environmental Sound Classification 50 (ESC-50) and Tampere University of Technology (TUT) Acoustic Scenes). Two teacher adaptation strategies are explored: classifier-head-only adaptation and full fine-tuning. The framework is further evaluated using various convolutional and transformer-based teacher models. Experimental results demonstrate consistent accuracy improvements across all datasets and settings, confirming the effectiveness and robustness of SSATKD in real-world sound classification tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
Ritu, Jarin
Mohammadi, Amirmohammad
Carreiro, Davelle
Van Dine, Alexandra
Peeples, Joshua
Sound
Audio and Speech Processing
While knowledge distillation has shown success in various audio tasks, its application to environmental sound classification often overlooks essential low-level audio texture features needed to capture local patterns in complex acoustic environments. To address this gap, the Structural and Statistical Audio Texture Knowledge Distillation (SSATKD) framework is proposed, which combines high-level contextual information with low-level structural and statistical audio textures extracted from intermediate layers. To evaluate its generalizability across diverse acoustic domains, SSATKD is tested on four datasets within the environmental sound classification domain, including two passive sonar datasets (DeepShip and Vessel Type Underwater Acoustic Data (VTUAD)) and two general environmental sound datasets (Environmental Sound Classification 50 (ESC-50) and Tampere University of Technology (TUT) Acoustic Scenes). Two teacher adaptation strategies are explored: classifier-head-only adaptation and full fine-tuning. The framework is further evaluated using various convolutional and transformer-based teacher models. Experimental results demonstrate consistent accuracy improvements across all datasets and settings, confirming the effectiveness and robustness of SSATKD in real-world sound classification tasks.
title Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.01921