On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Varshavsky-Hassid, Miri, Hirsch, Roy, Cohen, Regev, Golany, Tomer, Freedman, Daniel, Rivlin, Ehud
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914822603407360
author Varshavsky-Hassid, Miri
Hirsch, Roy
Cohen, Regev
Golany, Tomer
Freedman, Daniel
Rivlin, Ehud
author_facet Varshavsky-Hassid, Miri
Hirsch, Roy
Cohen, Regev
Golany, Tomer
Freedman, Daniel
Rivlin, Ehud
contents The incorporation of Denoising Diffusion Models (DDMs) in the Text-to-Speech (TTS) domain is rising, providing great value in synthesizing high quality speech. Although they exhibit impressive audio quality, the extent of their semantic capabilities is unknown, and controlling their synthesized speech's vocal properties remains a challenge. Inspired by recent advances in image synthesis, we explore the latent space of frozen TTS models, which is composed of the latent bottleneck activations of the DDM's denoiser. We identify that this space contains rich semantic information, and outline several novel methods for finding semantic directions within it, both supervised and unsupervised. We then demonstrate how these enable off-the-shelf audio editing, without any further training, architectural changes or data requirements. We present evidence of the semantic and acoustic qualities of the edited audio, and provide supplemental samples: https://latent-analysis-grad-tts.github.io/speech-samples/.
format Preprint
id arxiv_https___arxiv_org_abs_2402_12423
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
Varshavsky-Hassid, Miri
Hirsch, Roy
Cohen, Regev
Golany, Tomer
Freedman, Daniel
Rivlin, Ehud
Sound
Computation and Language
Machine Learning
Audio and Speech Processing
The incorporation of Denoising Diffusion Models (DDMs) in the Text-to-Speech (TTS) domain is rising, providing great value in synthesizing high quality speech. Although they exhibit impressive audio quality, the extent of their semantic capabilities is unknown, and controlling their synthesized speech's vocal properties remains a challenge. Inspired by recent advances in image synthesis, we explore the latent space of frozen TTS models, which is composed of the latent bottleneck activations of the DDM's denoiser. We identify that this space contains rich semantic information, and outline several novel methods for finding semantic directions within it, both supervised and unsupervised. We then demonstrate how these enable off-the-shelf audio editing, without any further training, architectural changes or data requirements. We present evidence of the semantic and acoustic qualities of the edited audio, and provide supplemental samples: https://latent-analysis-grad-tts.github.io/speech-samples/.
title On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
topic Sound
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2402.12423