Separate Anything You Describe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xubo, Kong, Qiuqiang, Zhao, Yan, Liu, Haohe, Yuan, Yi, Liu, Yuzhuo, Xia, Rui, Wang, Yuxuan, Plumbley, Mark D., Wang, Wenwu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910721208483840
author Liu, Xubo
Kong, Qiuqiang
Zhao, Yan
Liu, Haohe
Yuan, Yi
Liu, Yuzhuo
Xia, Rui
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
author_facet Liu, Xubo
Kong, Qiuqiang
Zhao, Yan
Liu, Haohe
Yuan, Yi
Liu, Yuzhuo
Xia, Rui
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
contents Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and scalable interface for digital audio applications. Recent works on LASS, despite attaining promising separation performance on specific sources (e.g., musical instruments, limited classes of audio events), are unable to separate audio concepts in the open domain. In this work, we introduce AudioSep, a foundation model for open-domain audio source separation with natural language queries. We train AudioSep on large-scale multimodal datasets and extensively evaluate its capabilities on numerous tasks including audio event separation, musical instrument separation, and speech enhancement. AudioSep demonstrates strong separation performance and impressive zero-shot generalization ability using audio captions or text labels as queries, substantially outperforming previous audio-queried and language-queried sound separation models. For reproducibility of this work, we will release the source code, evaluation benchmark and pre-trained model at: https://github.com/Audio-AGI/AudioSep.
format Preprint
id arxiv_https___arxiv_org_abs_2308_05037
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Separate Anything You Describe
Liu, Xubo
Kong, Qiuqiang
Zhao, Yan
Liu, Haohe
Yuan, Yi
Liu, Yuzhuo
Xia, Rui
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
Audio and Speech Processing
Artificial Intelligence
Multimedia
Sound
Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and scalable interface for digital audio applications. Recent works on LASS, despite attaining promising separation performance on specific sources (e.g., musical instruments, limited classes of audio events), are unable to separate audio concepts in the open domain. In this work, we introduce AudioSep, a foundation model for open-domain audio source separation with natural language queries. We train AudioSep on large-scale multimodal datasets and extensively evaluate its capabilities on numerous tasks including audio event separation, musical instrument separation, and speech enhancement. AudioSep demonstrates strong separation performance and impressive zero-shot generalization ability using audio captions or text labels as queries, substantially outperforming previous audio-queried and language-queried sound separation models. For reproducibility of this work, we will release the source code, evaluation benchmark and pre-trained model at: https://github.com/Audio-AGI/AudioSep.
title Separate Anything You Describe
topic Audio and Speech Processing
Artificial Intelligence
Multimedia
Sound
url https://arxiv.org/abs/2308.05037