Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tan, Ee-Leng, Yeow, Jun Wei, Peksi, Santi, Li, Haowen, Yang, Ziyi, Gan, Woon-Seng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911151210627072
author Tan, Ee-Leng
Yeow, Jun Wei
Peksi, Santi
Li, Haowen
Yang, Ziyi
Gan, Woon-Seng
author_facet Tan, Ee-Leng
Yeow, Jun Wei
Peksi, Santi
Li, Haowen
Yang, Ziyi
Gan, Woon-Seng
contents In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a teacher to a student model, aiming to achieve high performance with limited complexity. The proposed model is based on a CNN-GRU model and is trained solely using the TAU Urban Acoustic Scene 2022 Mobile development dataset, without utilizing any external datasets, except for MicIRP, which is used for device impulse response (DIR) augmentation. The proposed model has a memory usage of 114.2KB and requires 10.9M muliply-and-accumulate (MAC) operations. Using the development dataset, the proposed model achieved an accuracy of 60.25%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation
Tan, Ee-Leng
Yeow, Jun Wei
Peksi, Santi
Li, Haowen
Yang, Ziyi
Gan, Woon-Seng
Audio and Speech Processing
In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a teacher to a student model, aiming to achieve high performance with limited complexity. The proposed model is based on a CNN-GRU model and is trained solely using the TAU Urban Acoustic Scene 2022 Mobile development dataset, without utilizing any external datasets, except for MicIRP, which is used for device impulse response (DIR) augmentation. The proposed model has a memory usage of 114.2KB and requires 10.9M muliply-and-accumulate (MAC) operations. Using the development dataset, the proposed model achieved an accuracy of 60.25%.
title Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.09931