Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Turetzky, Arnon, Dekel, Avihu, Shabtay, Nimrod, Shechtman, Slava, Haws, David, Aronowitz, Hagai, Hoory, Ron, Adi, Yossi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908669871915008
author Turetzky, Arnon
Dekel, Avihu
Shabtay, Nimrod
Shechtman, Slava
Haws, David
Aronowitz, Hagai
Hoory, Ron
Adi, Yossi
author_facet Turetzky, Arnon
Dekel, Avihu
Shabtay, Nimrod
Shechtman, Slava
Haws, David
Aronowitz, Hagai
Hoory, Ron
Adi, Yossi
contents We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16048
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
Turetzky, Arnon
Dekel, Avihu
Shabtay, Nimrod
Shechtman, Slava
Haws, David
Aronowitz, Hagai
Hoory, Ron
Adi, Yossi
Audio and Speech Processing
We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio.
title Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
topic Audio and Speech Processing
url https://arxiv.org/abs/2410.16048