OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Xize, Zheng, Siqi, Wang, Zehan, Fang, Minghui, Zhang, Ziang, Huang, Rongjie, Ma, Ziyang, Ji, Shengpeng, Zuo, Jialong, Jin, Tao, Zhao, Zhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916457128919040
author Cheng, Xize
Zheng, Siqi
Wang, Zehan
Fang, Minghui
Zhang, Ziang
Huang, Rongjie
Ma, Ziyang
Ji, Shengpeng
Zuo, Jialong
Jin, Tao
Zhao, Zhou
author_facet Cheng, Xize
Zheng, Siqi
Wang, Zehan
Fang, Minghui
Zhang, Ziang
Huang, Rongjie
Ma, Ziyang
Ji, Shengpeng
Zuo, Jialong
Jin, Tao
Zhao, Zhou
contents The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21269
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
Cheng, Xize
Zheng, Siqi
Wang, Zehan
Fang, Minghui
Zhang, Ziang
Huang, Rongjie
Ma, Ziyang
Ji, Shengpeng
Zuo, Jialong
Jin, Tao
Zhao, Zhou
Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}.
title OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
topic Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2410.21269