MOSS-TTS Technical Report

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gong, Yitian, Jiang, Botian, Zhao, Yiwei, Yuan, Yucheng, Chen, Kuangwei, Jiang, Yaozhou, Chang, Cheng, Hong, Dong, Chen, Mingshu, Li, Ruixiao, Zhang, Yiyang, Gao, Yang, Chen, Hanfu, Chen, Ke, Wang, Songlin, Yang, Xiaogui, Zhang, Yuqian, Huang, Kexin, Lin, ZhengYuan, Yu, Kang, Chen, Ziqi, Wang, Jin, Fei, Zhaoye, Cheng, Qinyuan, Li, Shimin, Qiu, Xipeng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912975818850304
author Gong, Yitian
Jiang, Botian
Zhao, Yiwei
Yuan, Yucheng
Chen, Kuangwei
Jiang, Yaozhou
Chang, Cheng
Hong, Dong
Chen, Mingshu
Li, Ruixiao
Zhang, Yiyang
Gao, Yang
Chen, Hanfu
Chen, Ke
Wang, Songlin
Yang, Xiaogui
Zhang, Yuqian
Huang, Kexin
Lin, ZhengYuan
Yu, Kang
Chen, Ziqi
Wang, Jin
Fei, Zhaoye
Cheng, Qinyuan
Li, Shimin
Qiu, Xipeng
author_facet Gong, Yitian
Jiang, Botian
Zhao, Yiwei
Yuan, Yucheng
Chen, Kuangwei
Jiang, Yaozhou
Chang, Cheng
Hong, Dong
Chen, Mingshu
Li, Ruixiao
Zhang, Yiyang
Gao, Yang
Chen, Hanfu
Chen, Ke
Wang, Songlin
Yang, Xiaogui
Zhang, Yuqian
Huang, Kexin
Lin, ZhengYuan
Yu, Kang
Chen, Ziqi
Wang, Jin
Fei, Zhaoye
Cheng, Qinyuan
Li, Shimin
Qiu, Xipeng
contents This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18090
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MOSS-TTS Technical Report
Gong, Yitian
Jiang, Botian
Zhao, Yiwei
Yuan, Yucheng
Chen, Kuangwei
Jiang, Yaozhou
Chang, Cheng
Hong, Dong
Chen, Mingshu
Li, Ruixiao
Zhang, Yiyang
Gao, Yang
Chen, Hanfu
Chen, Ke
Wang, Songlin
Yang, Xiaogui
Zhang, Yuqian
Huang, Kexin
Lin, ZhengYuan
Yu, Kang
Chen, Ziqi
Wang, Jin
Fei, Zhaoye
Cheng, Qinyuan
Li, Shimin
Qiu, Xipeng
Sound
Artificial Intelligence
Computation and Language
This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models.
title MOSS-TTS Technical Report
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2603.18090