SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaofei, Thakker, Manthan, Chen, Zhuo, Kanda, Naoyuki, Eskimez, Sefik Emre, Chen, Sanyuan, Tang, Min, Liu, Shujie, Li, Jinyu, Yoshioka, Takuya
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917704882978816
author Wang, Xiaofei
Thakker, Manthan
Chen, Zhuo
Kanda, Naoyuki
Eskimez, Sefik Emre
Chen, Sanyuan
Tang, Min
Liu, Shujie
Li, Jinyu
Yoshioka, Takuya
author_facet Wang, Xiaofei
Thakker, Manthan
Chen, Zhuo
Kanda, Naoyuki
Eskimez, Sefik Emre
Chen, Sanyuan
Tang, Min
Liu, Shujie
Li, Jinyu
Yoshioka, Takuya
contents Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples.
format Preprint
id arxiv_https___arxiv_org_abs_2308_06873
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
Wang, Xiaofei
Thakker, Manthan
Chen, Zhuo
Kanda, Naoyuki
Eskimez, Sefik Emre
Chen, Sanyuan
Tang, Min
Liu, Shujie
Li, Jinyu
Yoshioka, Takuya
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling diverse audio-text speech generation tasks involving transforming input speech and processing audio captured in adverse acoustic conditions. This paper introduces SpeechX, a versatile speech generation model capable of zero-shot TTS and various speech transformation tasks, dealing with both clean and noisy signals. SpeechX combines neural codec language modeling with multi-task learning using task-dependent prompting, enabling unified and extensible modeling and providing a consistent way for leveraging textual input in speech enhancement and transformation tasks. Experimental results show SpeechX's efficacy in various tasks, including zero-shot TTS, noise suppression, target speaker extraction, speech removal, and speech editing with or without background noise, achieving comparable or superior performance to specialized models across tasks. See https://aka.ms/speechx for demo samples.
title SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2308.06873