Kimi-Audio Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909594439122944 |
|---|---|
| author | KimiTeam Ding, Ding Ju, Zeqian Leng, Yichong Liu, Songxiang Liu, Tong Shang, Zeyu Shen, Kai Song, Wei Tan, Xu Tang, Heyi Wang, Zhengtao Wei, Chu Xin, Yifei Xu, Xinran Yu, Jianwei Zhang, Yutao Zhou, Xinyu Charles, Y. Chen, Jun Chen, Yanru Du, Yulun He, Weiran Hu, Zhenxing Lai, Guokun Li, Qingcheng Liu, Yangyang Sun, Weidong Wang, Jianzhou Wang, Yuzhi Wu, Yuefeng Wu, Yuxin Yang, Dongchao Yang, Hao Yang, Ying Yang, Zhilin Yin, Aoxiong Yuan, Ruibin Zhang, Yutong Zhou, Zaida |
| author_facet | KimiTeam Ding, Ding Ju, Zeqian Leng, Yichong Liu, Songxiang Liu, Tong Shang, Zeyu Shen, Kai Song, Wei Tan, Xu Tang, Heyi Wang, Zhengtao Wei, Chu Xin, Yifei Xu, Xinran Yu, Jianwei Zhang, Yutao Zhou, Xinyu Charles, Y. Chen, Jun Chen, Yanru Du, Yulun He, Weiran Hu, Zhenxing Lai, Guokun Li, Qingcheng Liu, Yangyang Sun, Weidong Wang, Jianzhou Wang, Yuzhi Wu, Yuefeng Wu, Yuxin Yang, Dongchao Yang, Hao Yang, Ying Yang, Zhilin Yin, Aoxiong Yuan, Ruibin Zhang, Yutong Zhou, Zaida |
| contents | We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_18425 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Kimi-Audio Technical Report KimiTeam Ding, Ding Ju, Zeqian Leng, Yichong Liu, Songxiang Liu, Tong Shang, Zeyu Shen, Kai Song, Wei Tan, Xu Tang, Heyi Wang, Zhengtao Wei, Chu Xin, Yifei Xu, Xinran Yu, Jianwei Zhang, Yutao Zhou, Xinyu Charles, Y. Chen, Jun Chen, Yanru Du, Yulun He, Weiran Hu, Zhenxing Lai, Guokun Li, Qingcheng Liu, Yangyang Sun, Weidong Wang, Jianzhou Wang, Yuzhi Wu, Yuefeng Wu, Yuxin Yang, Dongchao Yang, Hao Yang, Ying Yang, Zhilin Yin, Aoxiong Yuan, Ruibin Zhang, Yutong Zhou, Zaida Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Multimedia Sound We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio. |
| title | Kimi-Audio Technical Report |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Machine Learning Multimedia Sound |
| url | https://arxiv.org/abs/2504.18425 |