Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Morocutti, Tobias, Schmid, Florian, Koutini, Khaled, Widmer, Gerhard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916652686245888
author Morocutti, Tobias
Schmid, Florian
Koutini, Khaled
Widmer, Gerhard
author_facet Morocutti, Tobias
Schmid, Florian
Koutini, Khaled
Widmer, Gerhard
contents Knowledge Distillation (KD) is a widespread technique for compressing the knowledge of large models into more compact and efficient models. KD has proved to be highly effective in building well-performing low-complexity Acoustic Scene Classification (ASC) systems and was used in all the top-ranked submissions to this task of the annual DCASE challenge in the past three years. There is extensive research available on establishing the KD process, designing efficient student models, and forming well-performing teacher ensembles. However, less research has been conducted on investigating which teacher model attributes are beneficial for low-complexity students. In this work, we try to close this gap by studying the effects on the student's performance when using different teacher network architectures, varying the teacher model size, training them with different device generalization methods, and applying different ensembling strategies. The results show that teacher model sizes, device generalization methods, the ensembling strategy and the ensemble size are key factors for a well-performing student network.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
Morocutti, Tobias
Schmid, Florian
Koutini, Khaled
Widmer, Gerhard
Sound
Machine Learning
Audio and Speech Processing
Knowledge Distillation (KD) is a widespread technique for compressing the knowledge of large models into more compact and efficient models. KD has proved to be highly effective in building well-performing low-complexity Acoustic Scene Classification (ASC) systems and was used in all the top-ranked submissions to this task of the annual DCASE challenge in the past three years. There is extensive research available on establishing the KD process, designing efficient student models, and forming well-performing teacher ensembles. However, less research has been conducted on investigating which teacher model attributes are beneficial for low-complexity students. In this work, we try to close this gap by studying the effects on the student's performance when using different teacher network architectures, varying the teacher model size, training them with different device generalization methods, and applying different ensembling strategies. The results show that teacher model sizes, device generalization methods, the ensembling strategy and the ensemble size are key factors for a well-performing student network.
title Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2503.11363