PromptSep: Generative Audio Separation via Multimodal Prompting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wen, Yutong, Chen, Ke, Seetharaman, Prem, Nieto, Oriol, Su, Jiaqi, Kumar, Rithesh, Kim, Minje, Smaragdis, Paris, Jin, Zeyu, Salamon, Justin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915602751291392
author Wen, Yutong
Chen, Ke
Seetharaman, Prem
Nieto, Oriol
Su, Jiaqi
Kumar, Rithesh
Kim, Minje
Smaragdis, Paris
Jin, Zeyu
Salamon, Justin
author_facet Wen, Yutong
Chen, Ke
Seetharaman, Prem
Nieto, Oriol
Su, Jiaqi
Kumar, Rithesh
Kim, Minje
Smaragdis, Paris
Jin, Zeyu
Salamon, Justin
contents Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PromptSep: Generative Audio Separation via Multimodal Prompting
Wen, Yutong
Chen, Ke
Seetharaman, Prem
Nieto, Oriol
Su, Jiaqi
Kumar, Rithesh
Kim, Minje
Smaragdis, Paris
Jin, Zeyu
Salamon, Justin
Sound
Audio and Speech Processing
Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.
title PromptSep: Generative Audio Separation via Multimodal Prompting
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2511.04623