Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908669871915008 |
|---|---|
| author | Turetzky, Arnon Dekel, Avihu Shabtay, Nimrod Shechtman, Slava Haws, David Aronowitz, Hagai Hoory, Ron Adi, Yossi |
| author_facet | Turetzky, Arnon Dekel, Avihu Shabtay, Nimrod Shechtman, Slava Haws, David Aronowitz, Hagai Hoory, Ron Adi, Yossi |
| contents | We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_16048 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion Turetzky, Arnon Dekel, Avihu Shabtay, Nimrod Shechtman, Slava Haws, David Aronowitz, Hagai Hoory, Ron Adi, Yossi Audio and Speech Processing We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio. |
| title | Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.16048 |