Kwai Keye-VL Technical Report
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913923160080384 |
|---|---|
| author | Kwai Keye Team Yang, Biao Wen, Bin Liu, Changyi Chu, Chenglong Song, Chengru Rao, Chongling Yi, Chuan Li, Da Zang, Dunju Yang, Fan Zhou, Guorui Peng, Hao Ding, Haojie Huang, Jiaming Cao, Jiangxia Chen, Jiankang Hua, Jingyun Ouyang, Jin Chen, Kaibing Jiang, Kaiyu Tang, Kaiyu Gai, Kun Zhang, Shengnan Mao, Siyang Huang, Sui Zhang, Tianke Gao, Tingting Chen, Wei Yuan, Wei Wu, Xiangyu Hu, Xiao Lu, Xingyu Zhou, Yang Zhang, Yi-Fan Yang, Yiping Chen, Yulong Wu, Zhenhua Li, Zhenyu Ling, Zhixin Li, Ziming Ma, Dehua Xu, Di Gao, Haixuan Li, Hang Guo, Jiawei Wang, Jing Ren, Lejian Wei, Muhao Wang, Qianqian Hu, Qigen Wang, Shiyao Yu, Tao Luo, Xinchen Li, Yan Liang, Yiming Hu, Yuhang Lu, Zeyi Yang, Zhuoran Zhang, Zixing |
| author_facet | Kwai Keye Team Yang, Biao Wen, Bin Liu, Changyi Chu, Chenglong Song, Chengru Rao, Chongling Yi, Chuan Li, Da Zang, Dunju Yang, Fan Zhou, Guorui Peng, Hao Ding, Haojie Huang, Jiaming Cao, Jiangxia Chen, Jiankang Hua, Jingyun Ouyang, Jin Chen, Kaibing Jiang, Kaiyu Tang, Kaiyu Gai, Kun Zhang, Shengnan Mao, Siyang Huang, Sui Zhang, Tianke Gao, Tingting Chen, Wei Yuan, Wei Wu, Xiangyu Hu, Xiao Lu, Xingyu Zhou, Yang Zhang, Yi-Fan Yang, Yiping Chen, Yulong Wu, Zhenhua Li, Zhenyu Ling, Zhixin Li, Ziming Ma, Dehua Xu, Di Gao, Haixuan Li, Hang Guo, Jiawei Wang, Jing Ren, Lejian Wei, Muhao Wang, Qianqian Hu, Qigen Wang, Shiyao Yu, Tao Luo, Xinchen Li, Yan Liang, Yiming Hu, Yuhang Lu, Zeyi Yang, Zhuoran Zhang, Zixing |
| contents | While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_01949 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Kwai Keye-VL Technical Report Kwai Keye Team Yang, Biao Wen, Bin Liu, Changyi Chu, Chenglong Song, Chengru Rao, Chongling Yi, Chuan Li, Da Zang, Dunju Yang, Fan Zhou, Guorui Peng, Hao Ding, Haojie Huang, Jiaming Cao, Jiangxia Chen, Jiankang Hua, Jingyun Ouyang, Jin Chen, Kaibing Jiang, Kaiyu Tang, Kaiyu Gai, Kun Zhang, Shengnan Mao, Siyang Huang, Sui Zhang, Tianke Gao, Tingting Chen, Wei Yuan, Wei Wu, Xiangyu Hu, Xiao Lu, Xingyu Zhou, Yang Zhang, Yi-Fan Yang, Yiping Chen, Yulong Wu, Zhenhua Li, Zhenyu Ling, Zhixin Li, Ziming Ma, Dehua Xu, Di Gao, Haixuan Li, Hang Guo, Jiawei Wang, Jing Ren, Lejian Wei, Muhao Wang, Qianqian Hu, Qigen Wang, Shiyao Yu, Tao Luo, Xinchen Li, Yan Liang, Yiming Hu, Yuhang Lu, Zeyi Yang, Zhuoran Zhang, Zixing Computer Vision and Pattern Recognition While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage. |
| title | Kwai Keye-VL Technical Report |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2507.01949 |