Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Haoning, Li, Zhaoqing, Chen, Youjun, Wang, Huimeng, Li, Guinan, Geng, Mengzhe, Deng, Chengxi, Liu, Xunying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913864403124224
author Xu, Haoning
Li, Zhaoqing
Chen, Youjun
Wang, Huimeng
Li, Guinan
Geng, Mengzhe
Deng, Chengxi
Liu, Xunying
author_facet Xu, Haoning
Li, Zhaoqing
Chen, Youjun
Wang, Huimeng
Li, Guinan
Geng, Mengzhe
Deng, Chengxi
Liu, Xunying
contents This paper presents a novel approach for speech foundation models compression that tightly integrates model pruning and parameter update into a single stage. Highly compact layer-level tied self-pinching gates each containing only a single learnable threshold are jointly trained with uncompressed models and used in fine-grained neuron level pruning. Experiments conducted on the LibriSpeech-100hr corpus suggest that our approach reduces the number of parameters of wav2vec2.0-base and HuBERT-large models by 65% and 60% respectively, while incurring no statistically significant word error rate (WER) increase on the test-clean dataset. Compared to previously published methods on the same task, our approach not only achieves the lowest WER of 7.05% on the test-clean dataset under a comparable model compression ratio of 4.26x, but also operates with at least 25% less model compression time.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
Xu, Haoning
Li, Zhaoqing
Chen, Youjun
Wang, Huimeng
Li, Guinan
Geng, Mengzhe
Deng, Chengxi
Liu, Xunying
Sound
Artificial Intelligence
Audio and Speech Processing
This paper presents a novel approach for speech foundation models compression that tightly integrates model pruning and parameter update into a single stage. Highly compact layer-level tied self-pinching gates each containing only a single learnable threshold are jointly trained with uncompressed models and used in fine-grained neuron level pruning. Experiments conducted on the LibriSpeech-100hr corpus suggest that our approach reduces the number of parameters of wav2vec2.0-base and HuBERT-large models by 65% and 60% respectively, while incurring no statistically significant word error rate (WER) increase on the test-clean dataset. Compared to previously published methods on the same task, our approach not only achieves the lowest WER of 7.05% on the test-clean dataset under a comparable model compression ratio of 4.26x, but also operates with at least 25% less model compression time.
title Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.22608