Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Jiangyu, Landini, Federico, Rohdin, Johan, Silnova, Anna, Diez, Mireia, Cernocky, Jan, Burget, Lukas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916767269388288
author Han, Jiangyu
Landini, Federico
Rohdin, Johan
Silnova, Anna
Diez, Mireia
Cernocky, Jan
Burget, Lukas
author_facet Han, Jiangyu
Landini, Federico
Rohdin, Johan
Silnova, Anna
Diez, Mireia
Cernocky, Jan
Burget, Lukas
contents Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource constrained scenarios. Previous studies have explored compression techniques, but usually for the price of degraded performance at high pruning ratios. In this work, we propose to compress SSL models through structured pruning by introducing knowledge distillation. Different from the existing works, we emphasize the importance of fine-tuning SSL models before pruning. Experiments on far-field single-channel AMI, AISHELL-4, and AliMeeting datasets show that our method can remove redundant parameters of WavLM Base+ and WavLM Large by up to 80% without any performance degradation. After pruning, the inference speeds on a single GPU for the Base+ and Large models are 4.0 and 2.6 times faster, respectively. Our source code is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Han, Jiangyu
Landini, Federico
Rohdin, Johan
Silnova, Anna
Diez, Mireia
Cernocky, Jan
Burget, Lukas
Audio and Speech Processing
Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource constrained scenarios. Previous studies have explored compression techniques, but usually for the price of degraded performance at high pruning ratios. In this work, we propose to compress SSL models through structured pruning by introducing knowledge distillation. Different from the existing works, we emphasize the importance of fine-tuning SSL models before pruning. Experiments on far-field single-channel AMI, AISHELL-4, and AliMeeting datasets show that our method can remove redundant parameters of WavLM Base+ and WavLM Large by up to 80% without any performance degradation. After pruning, the inference speeds on a single GPU for the Base+ and Large models are 4.0 and 2.6 times faster, respectively. Our source code is publicly available.
title Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
topic Audio and Speech Processing
url https://arxiv.org/abs/2505.24111