A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Watcharasupat, Karn N., Lerch, Alexander
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914923376803840
author Watcharasupat, Karn N.
Lerch, Alexander
author_facet Watcharasupat, Karn N.
Lerch, Alexander
contents Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems. Increasing stem support in these inflexible systems correspondingly requires increasing computational complexity, rendering extensions of these systems computationally infeasible for long-tail instruments. In this work, we propose Banquet, a system that allows source separation of multiple stems using just one decoder. A bandsplit source separation model is extended to work in a query-based setup in tandem with a music instrument recognition PaSST model. On the MoisesDB dataset, Banquet, at only 24.9 M trainable parameters, approached the performance level of the significantly more complex 6-stem Hybrid Transformer Demucs on VDBO stems and outperformed it on guitar and piano. The query-based setup allows for the separation of narrow instrument classes such as clean acoustic guitars, and can be successfully applied to the extraction of less common stems such as reeds and organs. Implementation is available at https://github.com/kwatcharasupat/query-bandit.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18747
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
Watcharasupat, Karn N.
Lerch, Alexander
Sound
Artificial Intelligence
Information Retrieval
Machine Learning
Audio and Speech Processing
Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems. Increasing stem support in these inflexible systems correspondingly requires increasing computational complexity, rendering extensions of these systems computationally infeasible for long-tail instruments. In this work, we propose Banquet, a system that allows source separation of multiple stems using just one decoder. A bandsplit source separation model is extended to work in a query-based setup in tandem with a music instrument recognition PaSST model. On the MoisesDB dataset, Banquet, at only 24.9 M trainable parameters, approached the performance level of the significantly more complex 6-stem Hybrid Transformer Demucs on VDBO stems and outperformed it on guitar and piano. The query-based setup allows for the separation of narrow instrument classes such as clean acoustic guitars, and can be successfully applied to the extraction of less common stems such as reeds and organs. Implementation is available at https://github.com/kwatcharasupat/query-bandit.
title A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
topic Sound
Artificial Intelligence
Information Retrieval
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2406.18747