VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913735056031744 |
|---|---|
| author | Du, Chenpeng Guo, Yiwei Wang, Hankun Yang, Yifan Niu, Zhikang Wang, Shuai Zhang, Hui Chen, Xie Yu, Kai |
| author_facet | Du, Chenpeng Guo, Yiwei Wang, Hankun Yang, Yifan Niu, Zhikang Wang, Shuai Zhang, Hui Chen, Xie Yu, Kai |
| contents | Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hallucination issues such as mispronunciation, word skipping and repeating. To address this limitation, we propose VALL-T, a generative Transducer model that introduces shifting relative position embeddings for input phoneme sequence, explicitly indicating the monotonic generation process while maintaining the architecture of decoder-only Transformer. Consequently, VALL-T retains the capability of prompt-based zero-shot adaptation and demonstrates better robustness against hallucinations with a relative reduction of 28.3% in the word error rate. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_14321 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech Du, Chenpeng Guo, Yiwei Wang, Hankun Yang, Yifan Niu, Zhikang Wang, Shuai Zhang, Hui Chen, Xie Yu, Kai Audio and Speech Processing Sound Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hallucination issues such as mispronunciation, word skipping and repeating. To address this limitation, we propose VALL-T, a generative Transducer model that introduces shifting relative position embeddings for input phoneme sequence, explicitly indicating the monotonic generation process while maintaining the architecture of decoder-only Transformer. Consequently, VALL-T retains the capability of prompt-based zero-shot adaptation and demonstrates better robustness against hallucinations with a relative reduction of 28.3% in the word error rate. |
| title | VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2401.14321 |