UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rong, Xiaobin, Yang, Leyan, Wang, Dahan, Hu, Yuxiang, Zhu, Changbao, Chen, Kai, Lu, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918316672548864
author Rong, Xiaobin
Yang, Leyan
Wang, Dahan
Hu, Yuxiang
Zhu, Changbao
Chen, Kai
Lu, Jing
author_facet Rong, Xiaobin
Yang, Leyan
Wang, Dahan
Hu, Yuxiang
Zhu, Changbao
Chen, Kai
Lu, Jing
contents Lightweight models are essential for real-time speech enhancement applications. In recent years, there has been a growing trend toward developing increasingly compact models for speech enhancement. In this paper, we propose an Ultra-Lightweight U-net optimized by Network Architecture Search (UL-UNAS), which is suitable for implementation in low-footprint devices. Firstly, we explore the application of various efficient convolutional blocks within the U-Net framework to identify the most promising candidates. Secondly, we introduce two boosting components to enhance the capacity of these convolutional blocks: a novel activation function named affine PReLU and a causal time-frequency attention module. Furthermore, we leverage neural architecture search to discover an optimal architecture within our carefully designed search space. By integrating the above strategies, UL-UNAS not only significantly outperforms the latest ultra-lightweight models with the same or lower computational complexity, but also delivers competitive performance compared to recent baseline models that require substantially higher computational resources. Source code and audio demos are available at https://github.com/Xiaobin-Rong/ul-unas.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
Rong, Xiaobin
Yang, Leyan
Wang, Dahan
Hu, Yuxiang
Zhu, Changbao
Chen, Kai
Lu, Jing
Audio and Speech Processing
Lightweight models are essential for real-time speech enhancement applications. In recent years, there has been a growing trend toward developing increasingly compact models for speech enhancement. In this paper, we propose an Ultra-Lightweight U-net optimized by Network Architecture Search (UL-UNAS), which is suitable for implementation in low-footprint devices. Firstly, we explore the application of various efficient convolutional blocks within the U-Net framework to identify the most promising candidates. Secondly, we introduce two boosting components to enhance the capacity of these convolutional blocks: a novel activation function named affine PReLU and a causal time-frequency attention module. Furthermore, we leverage neural architecture search to discover an optimal architecture within our carefully designed search space. By integrating the above strategies, UL-UNAS not only significantly outperforms the latest ultra-lightweight models with the same or lower computational complexity, but also delivers competitive performance compared to recent baseline models that require substantially higher computational resources. Source code and audio demos are available at https://github.com/Xiaobin-Rong/ul-unas.
title UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
topic Audio and Speech Processing
url https://arxiv.org/abs/2503.00340