SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dong, Zhang, Xin, Zhan, Jun, Li, Shimin, Zhou, Yaqian, Qiu, Xipeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913208602722304
author Zhang, Dong
Zhang, Xin
Zhan, Jun
Li, Shimin
Zhou, Yaqian
Qiu, Xipeng
author_facet Zhang, Dong
Zhang, Xin
Zhan, Jun
Li, Shimin
Zhou, Yaqian
Qiu, Xipeng
contents Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing information modeling process is encumbered by certain redundancies, leading to inefficiencies in speech generation. We propose Chain-of-Information Generation (CoIG), a method for decoupling semantic and perceptual information in large-scale speech generation. Building on this, we develop SpeechGPT-Gen, an 8-billion-parameter SLLM efficient in semantic and perceptual information modeling. It comprises an autoregressive model based on LLM for semantic information modeling and a non-autoregressive model employing flow matching for perceptual information modeling. Additionally, we introduce the novel approach of infusing semantic information into the prior distribution to enhance the efficiency of flow matching. Extensive experimental results demonstrate that SpeechGPT-Gen markedly excels in zero-shot text-to-speech, zero-shot voice conversion, and speech-to-speech dialogue, underscoring CoIG's remarkable proficiency in capturing and modeling speech's semantic and perceptual dimensions. Code and models are available at https://github.com/0nutation/SpeechGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2401_13527
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
Zhang, Dong
Zhang, Xin
Zhan, Jun
Li, Shimin
Zhou, Yaqian
Qiu, Xipeng
Computation and Language
Sound
Audio and Speech Processing
Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers. However, the prevailing information modeling process is encumbered by certain redundancies, leading to inefficiencies in speech generation. We propose Chain-of-Information Generation (CoIG), a method for decoupling semantic and perceptual information in large-scale speech generation. Building on this, we develop SpeechGPT-Gen, an 8-billion-parameter SLLM efficient in semantic and perceptual information modeling. It comprises an autoregressive model based on LLM for semantic information modeling and a non-autoregressive model employing flow matching for perceptual information modeling. Additionally, we introduce the novel approach of infusing semantic information into the prior distribution to enhance the efficiency of flow matching. Extensive experimental results demonstrate that SpeechGPT-Gen markedly excels in zero-shot text-to-speech, zero-shot voice conversion, and speech-to-speech dialogue, underscoring CoIG's remarkable proficiency in capturing and modeling speech's semantic and perceptual dimensions. Code and models are available at https://github.com/0nutation/SpeechGPT.
title SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.13527