MOSS-TTS Technical Report
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912975818850304 |
|---|---|
| author | Gong, Yitian Jiang, Botian Zhao, Yiwei Yuan, Yucheng Chen, Kuangwei Jiang, Yaozhou Chang, Cheng Hong, Dong Chen, Mingshu Li, Ruixiao Zhang, Yiyang Gao, Yang Chen, Hanfu Chen, Ke Wang, Songlin Yang, Xiaogui Zhang, Yuqian Huang, Kexin Lin, ZhengYuan Yu, Kang Chen, Ziqi Wang, Jin Fei, Zhaoye Cheng, Qinyuan Li, Shimin Qiu, Xipeng |
| author_facet | Gong, Yitian Jiang, Botian Zhao, Yiwei Yuan, Yucheng Chen, Kuangwei Jiang, Yaozhou Chang, Cheng Hong, Dong Chen, Mingshu Li, Ruixiao Zhang, Yiyang Gao, Yang Chen, Hanfu Chen, Ke Wang, Songlin Yang, Xiaogui Zhang, Yuqian Huang, Kexin Lin, ZhengYuan Yu, Kang Chen, Ziqi Wang, Jin Fei, Zhaoye Cheng, Qinyuan Li, Shimin Qiu, Xipeng |
| contents | This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_18090 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MOSS-TTS Technical Report Gong, Yitian Jiang, Botian Zhao, Yiwei Yuan, Yucheng Chen, Kuangwei Jiang, Yaozhou Chang, Cheng Hong, Dong Chen, Mingshu Li, Ruixiao Zhang, Yiyang Gao, Yang Chen, Hanfu Chen, Ke Wang, Songlin Yang, Xiaogui Zhang, Yuqian Huang, Kexin Lin, ZhengYuan Yu, Kang Chen, Ziqi Wang, Jin Fei, Zhaoye Cheng, Qinyuan Li, Shimin Qiu, Xipeng Sound Artificial Intelligence Computation and Language This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models. |
| title | MOSS-TTS Technical Report |
| topic | Sound Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2603.18090 |