Towards Practical Real-Time Low-Latency Music Source Separation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Junyu, Liu, Jie, Pan, Tianrui, Tang, Jie, Wu, Gangshan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917085425172480
author Wu, Junyu
Liu, Jie
Pan, Tianrui
Tang, Jie
Wu, Gangshan
author_facet Wu, Junyu
Liu, Jie
Pan, Tianrui
Tang, Jie
Wu, Gangshan
contents In recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Practical Real-Time Low-Latency Music Source Separation
Wu, Junyu
Liu, Jie
Pan, Tianrui
Tang, Jie
Wu, Gangshan
Sound
Multimedia
In recent years, significant progress has been made in the field of deep learning for music demixing. However, there has been limited attention on real-time, low-latency music demixing, which holds potential for various applications, such as hearing aids, audio stream remixing, and live performances. Additionally, a notable tendency has emerged towards the development of larger models, limiting their applicability in certain scenarios. In this paper, we introduce a lightweight real-time low-latency model called Real-Time Single-Path TFC-TDF UNET (RT-STT), which is based on the Dual-Path TFC-TDF UNET (DTTNet). In RT-STT, we propose a feature fusion technique based on channel expansion. We also demonstrate the superiority of single-path modeling over dual-path modeling in real-time models. Moreover, we investigate the method of quantization to further reduce inference time. RT-STT exhibits superior performance with significantly fewer parameters and shorter inference times compared to state-of-the-art models.
title Towards Practical Real-Time Low-Latency Music Source Separation
topic Sound
Multimedia
url https://arxiv.org/abs/2511.13146