Qwen3-TTS Technical Report

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Hangrui, Zhu, Xinfa, He, Ting, Guo, Dake, Zhang, Bin, Wang, Xiong, Guo, Zhifang, Jiang, Ziyue, Hao, Hongkun, Guo, Zishan, Zhang, Xinyu, Zhang, Pei, Yang, Baosong, Xu, Jin, Zhou, Jingren, Lin, Junyang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912840479145984
author Hu, Hangrui
Zhu, Xinfa
He, Ting
Guo, Dake
Zhang, Bin
Wang, Xiong
Guo, Zhifang
Jiang, Ziyue
Hao, Hongkun
Guo, Zishan
Zhang, Xinyu
Zhang, Pei
Yang, Baosong
Xu, Jin
Zhou, Jingren
Lin, Junyang
author_facet Hu, Hangrui
Zhu, Xinfa
He, Ting
Guo, Dake
Zhang, Bin
Wang, Xiong
Guo, Zhifang
Jiang, Ziyue
Hao, Hongkun
Guo, Zishan
Zhang, Xinyu
Zhang, Pei
Yang, Baosong
Xu, Jin
Zhou, Jingren
Lin, Junyang
contents In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15621
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Qwen3-TTS Technical Report
Hu, Hangrui
Zhu, Xinfa
He, Ting
Guo, Dake
Zhang, Bin
Wang, Xiong
Guo, Zhifang
Jiang, Ziyue
Hao, Hongkun
Guo, Zishan
Zhang, Xinyu
Zhang, Pei
Yang, Baosong
Xu, Jin
Zhou, Jingren
Lin, Junyang
Sound
Computation and Language
Audio and Speech Processing
In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control, allowing both the creation of entirely novel voices and fine-grained manipulation over the output speech. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS adopts a dual-track LM architecture for real-time synthesis, coupled with two speech tokenizers: 1) Qwen-TTS-Tokenizer-25Hz is a single-codebook codec emphasizing semantic content, which offers seamlessly integration with Qwen-Audio and enables streaming waveform reconstruction via a block-wise DiT. 2) Qwen-TTS-Tokenizer-12Hz achieves extreme bitrate reduction and ultra-low-latency streaming, enabling immediate first-packet emission ($97\,\mathrm{ms}$) through its 12.5 Hz, 16-layer multi-codebook design and a lightweight causal ConvNet. Extensive experiments indicate state-of-the-art performance across diverse objective and subjective benchmark (e.g., TTS multilingual test set, InstructTTSEval, and our long speech test set). To facilitate community research and development, we release both tokenizers and models under the Apache 2.0 license.
title Qwen3-TTS Technical Report
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2601.15621