Improving Music Source Separation with Diffusion and Consistency Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karchkhadze, Tornike, Izadi, Mohammad Rasool, Zhang, Shuo, Dubnov, Shlomo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918466853797888
author Karchkhadze, Tornike
Izadi, Mohammad Rasool
Zhang, Shuo
Dubnov, Shlomo
author_facet Karchkhadze, Tornike
Izadi, Mohammad Rasool
Zhang, Shuo
Dubnov, Shlomo
contents In this work, we propose an approach to music source separation that uses a generative diffusion model as a last-stage refinement on top of a deterministic separator, progressively enhancing the separated sources through iterative denoising. While the diffusion refinement yields measurable quality gains, it requires iterative steps at inference, increasing computational cost. To speed up the inference process, we apply consistency distillation, reducing inference to a single step while maintaining quality; with two or more steps, the distilled model even surpasses the diffusion-based approach. Crucially, our method is architecture-agnostic: we demonstrate state-of-the-art results when applied to both a custom U-Net-based separator on Slakh2100 and the state-of-the-art BS-RoFormer model on MUSDB18, showing that the refinement generalizes across backbone architectures. Sound examples are available at: https://consistency-separation.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06965
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Music Source Separation with Diffusion and Consistency Refinement
Karchkhadze, Tornike
Izadi, Mohammad Rasool
Zhang, Shuo
Dubnov, Shlomo
Sound
Audio and Speech Processing
In this work, we propose an approach to music source separation that uses a generative diffusion model as a last-stage refinement on top of a deterministic separator, progressively enhancing the separated sources through iterative denoising. While the diffusion refinement yields measurable quality gains, it requires iterative steps at inference, increasing computational cost. To speed up the inference process, we apply consistency distillation, reducing inference to a single step while maintaining quality; with two or more steps, the distilled model even surpasses the diffusion-based approach. Crucially, our method is architecture-agnostic: we demonstrate state-of-the-art results when applied to both a custom U-Net-based separator on Slakh2100 and the state-of-the-art BS-RoFormer model on MUSDB18, showing that the refinement generalizes across backbone architectures. Sound examples are available at: https://consistency-separation.github.io/.
title Improving Music Source Separation with Diffusion and Consistency Refinement
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.06965