DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Geonyoung, Han, Geonhee, Seo, Paul Hongsuck
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908420700897280
author Lee, Geonyoung
Han, Geonhee
Seo, Paul Hongsuck
author_facet Lee, Geonyoung
Han, Geonhee
Seo, Paul Hongsuck
contents Language-queried Audio Source Separation (LASS) enables open-vocabulary sound separation via natural language queries. While existing methods rely on task-specific training, we explore whether pretrained diffusion models, originally designed for audio generation, can inherently perform separation without further training. In this study, we introduce a training-free framework leveraging generative priors for zero-shot LASS. Analyzing naive adaptations, we identify key limitations arising from modality-specific challenges. To address these issues, we propose Diffusion-Guided Mask Optimization (DGMO), a test-time optimization framework that refines spectrogram masks for precise, input-aligned separation. Our approach effectively repurposes pretrained diffusion models for source separation, achieving competitive performance without task-specific supervision. This work expands the application of diffusion models beyond generation, establishing a new paradigm for zero-shot audio separation. The code is available at: https://wltschmrz.github.io/DGMO/
format Preprint
id arxiv_https___arxiv_org_abs_2506_02858
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
Lee, Geonyoung
Han, Geonhee
Seo, Paul Hongsuck
Audio and Speech Processing
Artificial Intelligence
Sound
Language-queried Audio Source Separation (LASS) enables open-vocabulary sound separation via natural language queries. While existing methods rely on task-specific training, we explore whether pretrained diffusion models, originally designed for audio generation, can inherently perform separation without further training. In this study, we introduce a training-free framework leveraging generative priors for zero-shot LASS. Analyzing naive adaptations, we identify key limitations arising from modality-specific challenges. To address these issues, we propose Diffusion-Guided Mask Optimization (DGMO), a test-time optimization framework that refines spectrogram masks for precise, input-aligned separation. Our approach effectively repurposes pretrained diffusion models for source separation, achieving competitive performance without task-specific supervision. This work expands the application of diffusion models beyond generation, establishing a new paradigm for zero-shot audio separation. The code is available at: https://wltschmrz.github.io/DGMO/
title DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2506.02858