VibeVoice Technical Report
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914006957031424 |
|---|---|
| author | Peng, Zhiliang Yu, Jianwei Wang, Wenhui Chang, Yaoyao Sun, Yutao Dong, Li Zhu, Yi Xu, Weijiang Bao, Hangbo Wang, Zehua Huang, Shaohan Xia, Yan Wei, Furu |
| author_facet | Peng, Zhiliang Yu, Jianwei Wang, Wenhui Chang, Yaoyao Sun, Yutao Dong, Li Zhu, Yi Xu, Weijiang Bao, Hangbo Wang, Zehua Huang, Shaohan Xia, Yan Wei, Furu |
| contents | This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_19205 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VibeVoice Technical Report Peng, Zhiliang Yu, Jianwei Wang, Wenhui Chang, Yaoyao Sun, Yutao Dong, Li Zhu, Yi Xu, Weijiang Bao, Hangbo Wang, Zehua Huang, Shaohan Xia, Yan Wei, Furu Computation and Language Artificial Intelligence Sound Audio and Speech Processing This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models. |
| title | VibeVoice Technical Report |
| topic | Computation and Language Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.19205 |