_version_ 1866909594439122944
author KimiTeam
Ding, Ding
Ju, Zeqian
Leng, Yichong
Liu, Songxiang
Liu, Tong
Shang, Zeyu
Shen, Kai
Song, Wei
Tan, Xu
Tang, Heyi
Wang, Zhengtao
Wei, Chu
Xin, Yifei
Xu, Xinran
Yu, Jianwei
Zhang, Yutao
Zhou, Xinyu
Charles, Y.
Chen, Jun
Chen, Yanru
Du, Yulun
He, Weiran
Hu, Zhenxing
Lai, Guokun
Li, Qingcheng
Liu, Yangyang
Sun, Weidong
Wang, Jianzhou
Wang, Yuzhi
Wu, Yuefeng
Wu, Yuxin
Yang, Dongchao
Yang, Hao
Yang, Ying
Yang, Zhilin
Yin, Aoxiong
Yuan, Ruibin
Zhang, Yutong
Zhou, Zaida
author_facet KimiTeam
Ding, Ding
Ju, Zeqian
Leng, Yichong
Liu, Songxiang
Liu, Tong
Shang, Zeyu
Shen, Kai
Song, Wei
Tan, Xu
Tang, Heyi
Wang, Zhengtao
Wei, Chu
Xin, Yifei
Xu, Xinran
Yu, Jianwei
Zhang, Yutao
Zhou, Xinyu
Charles, Y.
Chen, Jun
Chen, Yanru
Du, Yulun
He, Weiran
Hu, Zhenxing
Lai, Guokun
Li, Qingcheng
Liu, Yangyang
Sun, Weidong
Wang, Jianzhou
Wang, Yuzhi
Wu, Yuefeng
Wu, Yuxin
Yang, Dongchao
Yang, Hao
Yang, Ying
Yang, Zhilin
Yin, Aoxiong
Yuan, Ruibin
Zhang, Yutong
Zhou, Zaida
contents We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18425
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Kimi-Audio Technical Report
KimiTeam
Ding, Ding
Ju, Zeqian
Leng, Yichong
Liu, Songxiang
Liu, Tong
Shang, Zeyu
Shen, Kai
Song, Wei
Tan, Xu
Tang, Heyi
Wang, Zhengtao
Wei, Chu
Xin, Yifei
Xu, Xinran
Yu, Jianwei
Zhang, Yutao
Zhou, Xinyu
Charles, Y.
Chen, Jun
Chen, Yanru
Du, Yulun
He, Weiran
Hu, Zhenxing
Lai, Guokun
Li, Qingcheng
Liu, Yangyang
Sun, Weidong
Wang, Jianzhou
Wang, Yuzhi
Wu, Yuefeng
Wu, Yuxin
Yang, Dongchao
Yang, Hao
Yang, Ying
Yang, Zhilin
Yin, Aoxiong
Yuan, Ruibin
Zhang, Yutong
Zhou, Zaida
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Sound
We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.
title Kimi-Audio Technical Report
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Sound
url https://arxiv.org/abs/2504.18425