_version_ 1866913923160080384
author Kwai Keye Team
Yang, Biao
Wen, Bin
Liu, Changyi
Chu, Chenglong
Song, Chengru
Rao, Chongling
Yi, Chuan
Li, Da
Zang, Dunju
Yang, Fan
Zhou, Guorui
Peng, Hao
Ding, Haojie
Huang, Jiaming
Cao, Jiangxia
Chen, Jiankang
Hua, Jingyun
Ouyang, Jin
Chen, Kaibing
Jiang, Kaiyu
Tang, Kaiyu
Gai, Kun
Zhang, Shengnan
Mao, Siyang
Huang, Sui
Zhang, Tianke
Gao, Tingting
Chen, Wei
Yuan, Wei
Wu, Xiangyu
Hu, Xiao
Lu, Xingyu
Zhou, Yang
Zhang, Yi-Fan
Yang, Yiping
Chen, Yulong
Wu, Zhenhua
Li, Zhenyu
Ling, Zhixin
Li, Ziming
Ma, Dehua
Xu, Di
Gao, Haixuan
Li, Hang
Guo, Jiawei
Wang, Jing
Ren, Lejian
Wei, Muhao
Wang, Qianqian
Hu, Qigen
Wang, Shiyao
Yu, Tao
Luo, Xinchen
Li, Yan
Liang, Yiming
Hu, Yuhang
Lu, Zeyi
Yang, Zhuoran
Zhang, Zixing
author_facet Kwai Keye Team
Yang, Biao
Wen, Bin
Liu, Changyi
Chu, Chenglong
Song, Chengru
Rao, Chongling
Yi, Chuan
Li, Da
Zang, Dunju
Yang, Fan
Zhou, Guorui
Peng, Hao
Ding, Haojie
Huang, Jiaming
Cao, Jiangxia
Chen, Jiankang
Hua, Jingyun
Ouyang, Jin
Chen, Kaibing
Jiang, Kaiyu
Tang, Kaiyu
Gai, Kun
Zhang, Shengnan
Mao, Siyang
Huang, Sui
Zhang, Tianke
Gao, Tingting
Chen, Wei
Yuan, Wei
Wu, Xiangyu
Hu, Xiao
Lu, Xingyu
Zhou, Yang
Zhang, Yi-Fan
Yang, Yiping
Chen, Yulong
Wu, Zhenhua
Li, Zhenyu
Ling, Zhixin
Li, Ziming
Ma, Dehua
Xu, Di
Gao, Haixuan
Li, Hang
Guo, Jiawei
Wang, Jing
Ren, Lejian
Wei, Muhao
Wang, Qianqian
Hu, Qigen
Wang, Shiyao
Yu, Tao
Luo, Xinchen
Li, Yan
Liang, Yiming
Hu, Yuhang
Lu, Zeyi
Yang, Zhuoran
Zhang, Zixing
contents While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Kwai Keye-VL Technical Report
Kwai Keye Team
Yang, Biao
Wen, Bin
Liu, Changyi
Chu, Chenglong
Song, Chengru
Rao, Chongling
Yi, Chuan
Li, Da
Zang, Dunju
Yang, Fan
Zhou, Guorui
Peng, Hao
Ding, Haojie
Huang, Jiaming
Cao, Jiangxia
Chen, Jiankang
Hua, Jingyun
Ouyang, Jin
Chen, Kaibing
Jiang, Kaiyu
Tang, Kaiyu
Gai, Kun
Zhang, Shengnan
Mao, Siyang
Huang, Sui
Zhang, Tianke
Gao, Tingting
Chen, Wei
Yuan, Wei
Wu, Xiangyu
Hu, Xiao
Lu, Xingyu
Zhou, Yang
Zhang, Yi-Fan
Yang, Yiping
Chen, Yulong
Wu, Zhenhua
Li, Zhenyu
Ling, Zhixin
Li, Ziming
Ma, Dehua
Xu, Di
Gao, Haixuan
Li, Hang
Guo, Jiawei
Wang, Jing
Ren, Lejian
Wei, Muhao
Wang, Qianqian
Hu, Qigen
Wang, Shiyao
Yu, Tao
Luo, Xinchen
Li, Yan
Liang, Yiming
Hu, Yuhang
Lu, Zeyi
Yang, Zhuoran
Zhang, Zixing
Computer Vision and Pattern Recognition
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage.
title Kwai Keye-VL Technical Report
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.01949