Dynamic Slimmable Networks for Efficient Speech Separation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elminshawi, Mohamed, Chetupalli, Srikanth Raj, Habets, Emanuël A. P.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918086502776832
author Elminshawi, Mohamed
Chetupalli, Srikanth Raj
Habets, Emanuël A. P.
author_facet Elminshawi, Mohamed
Chetupalli, Srikanth Raj
Habets, Emanuël A. P.
contents Recent progress in speech separation has been largely driven by advances in deep neural networks, yet their high computational and memory requirements hinder deployment on resource-constrained devices. A significant inefficiency in conventional systems arises from using static network architectures that maintain constant computational complexity across all input segments, regardless of their characteristics. This approach is sub-optimal for simpler segments that do not require intensive processing, such as silence or non-overlapping speech. To address this limitation, we propose a dynamic slimmable network (DSN) for speech separation that adaptively adjusts its computational complexity based on the input signal. The DSN combines a slimmable network, which can operate at different network widths, with a lightweight gating module that dynamically determines the required width by analyzing the local input characteristics. To balance performance and efficiency, we introduce a signal-dependent complexity loss that penalizes unnecessary computation based on segmental reconstruction error. Experiments on clean and noisy two-speaker mixtures from the WSJ0-2mix and WHAM! datasets show that the DSN achieves a better performance-efficiency trade-off than individually trained static networks of different sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Slimmable Networks for Efficient Speech Separation
Elminshawi, Mohamed
Chetupalli, Srikanth Raj
Habets, Emanuël A. P.
Audio and Speech Processing
Recent progress in speech separation has been largely driven by advances in deep neural networks, yet their high computational and memory requirements hinder deployment on resource-constrained devices. A significant inefficiency in conventional systems arises from using static network architectures that maintain constant computational complexity across all input segments, regardless of their characteristics. This approach is sub-optimal for simpler segments that do not require intensive processing, such as silence or non-overlapping speech. To address this limitation, we propose a dynamic slimmable network (DSN) for speech separation that adaptively adjusts its computational complexity based on the input signal. The DSN combines a slimmable network, which can operate at different network widths, with a lightweight gating module that dynamically determines the required width by analyzing the local input characteristics. To balance performance and efficiency, we introduce a signal-dependent complexity loss that penalizes unnecessary computation based on segmental reconstruction error. Experiments on clean and noisy two-speaker mixtures from the WSJ0-2mix and WHAM! datasets show that the DSN achieves a better performance-efficiency trade-off than individually trained static networks of different sizes.
title Dynamic Slimmable Networks for Efficient Speech Separation
topic Audio and Speech Processing
url https://arxiv.org/abs/2507.06179