Generative Pre-training for Speech with Flow Matching

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Alexander H., Le, Matt, Vyas, Apoorv, Shi, Bowen, Tjandra, Andros, Hsu, Wei-Ning
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909149705535488
author Liu, Alexander H.
Le, Matt
Vyas, Apoorv
Shi, Bowen
Tjandra, Andros
Hsu, Wei-Ning
author_facet Liu, Alexander H.
Le, Matt
Vyas, Apoorv
Shi, Bowen
Tjandra, Andros
Hsu, Wei-Ning
contents Generative models have gained more and more attention in recent years for their remarkable success in tasks that required estimating and sampling data distribution to generate high-fidelity synthetic data. In speech, text-to-speech synthesis and neural vocoder are good examples where generative models have shined. While generative models have been applied to different applications in speech, there exists no general-purpose generative model that models speech directly. In this work, we take a step toward this direction by showing a single pre-trained generative model can be adapted to different downstream tasks with strong performance. Specifically, we pre-trained a generative model, named SpeechFlow, on 60k hours of untranscribed speech with Flow Matching and masked conditions. Experiment results show the pre-trained generative model can be fine-tuned with task-specific data to match or surpass existing expert models on speech enhancement, separation, and synthesis. Our work suggested a foundational model for generation tasks in speech can be built with generative pre-training.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16338
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generative Pre-training for Speech with Flow Matching
Liu, Alexander H.
Le, Matt
Vyas, Apoorv
Shi, Bowen
Tjandra, Andros
Hsu, Wei-Ning
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
Generative models have gained more and more attention in recent years for their remarkable success in tasks that required estimating and sampling data distribution to generate high-fidelity synthetic data. In speech, text-to-speech synthesis and neural vocoder are good examples where generative models have shined. While generative models have been applied to different applications in speech, there exists no general-purpose generative model that models speech directly. In this work, we take a step toward this direction by showing a single pre-trained generative model can be adapted to different downstream tasks with strong performance. Specifically, we pre-trained a generative model, named SpeechFlow, on 60k hours of untranscribed speech with Flow Matching and masked conditions. Experiment results show the pre-trained generative model can be fine-tuned with task-specific data to match or surpass existing expert models on speech enhancement, separation, and synthesis. Our work suggested a foundational model for generation tasks in speech can be built with generative pre-training.
title Generative Pre-training for Speech with Flow Matching
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2310.16338