VibeVoice Technical Report

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Zhiliang, Yu, Jianwei, Wang, Wenhui, Chang, Yaoyao, Sun, Yutao, Dong, Li, Zhu, Yi, Xu, Weijiang, Bao, Hangbo, Wang, Zehua, Huang, Shaohan, Xia, Yan, Wei, Furu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914006957031424
author Peng, Zhiliang
Yu, Jianwei
Wang, Wenhui
Chang, Yaoyao
Sun, Yutao
Dong, Li
Zhu, Yi
Xu, Weijiang
Bao, Hangbo
Wang, Zehua
Huang, Shaohan
Xia, Yan
Wei, Furu
author_facet Peng, Zhiliang
Yu, Jianwei
Wang, Wenhui
Chang, Yaoyao
Sun, Yutao
Dong, Li
Zhu, Yi
Xu, Weijiang
Bao, Hangbo
Wang, Zehua
Huang, Shaohan
Xia, Yan
Wei, Furu
contents This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19205
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VibeVoice Technical Report
Peng, Zhiliang
Yu, Jianwei
Wang, Wenhui
Chang, Yaoyao
Sun, Yutao
Dong, Li
Zhu, Yi
Xu, Weijiang
Bao, Hangbo
Wang, Zehua
Huang, Shaohan
Xia, Yan
Wei, Furu
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.
title VibeVoice Technical Report
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.19205