SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917704882978816 |
|---|---|
| author | Wang, Xiaofei Thakker, Manthan Chen, Zhuo Kanda, Naoyuki Eskimez, Sefik Emre Chen, Sanyuan Tang, Min Liu, Shujie Li, Jinyu Yoshioka, Takuya |
| author_facet | Wang, Xiaofei Thakker, Manthan Chen, Zhuo Kanda, Naoyuki Eskimez, Sefik Emre Chen, Sanyuan Tang, Min Liu, Shujie Li, Jinyu Yoshioka, Takuya |
| contents | Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2308_06873 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SpeechX: Neural Codec Language Model as a Versatile Speech Transformer Wang, Xiaofei Thakker, Manthan Chen, Zhuo Kanda, Naoyuki Eskimez, Sefik Emre Chen, Sanyuan Tang, Min Liu, Shujie Li, Jinyu Yoshioka, Takuya Audio and Speech Processing Computation and Language Machine Learning Sound Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples. |
| title | SpeechX: Neural Codec Language Model as a Versatile Speech Transformer |
| topic | Audio and Speech Processing Computation and Language Machine Learning Sound |
| url | https://arxiv.org/abs/2308.06873 |