Kling-Omni Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909969343840256 |
|---|---|
| author | Kling Team Chen, Jialu Ci, Yuanzheng Du, Xiangyu Feng, Zipeng Gai, Kun Guo, Sainan Han, Feng He, Jingbin He, Kang Hu, Xiao Hu, Xiaohua Jiang, Boyuan Kong, Fangyuan Li, Hang Li, Jie Li, Qingyu Li, Shen Li, Xiaohan Li, Yan Liang, Jiajun Liao, Borui Liao, Yiqiao Lin, Weihong Liu, Quande Liu, Xiaokun Liu, Yilun Liu, Yuliang Lu, Shun Mao, Hangyu Mao, Yunyao Ouyang, Haodong Qin, Wenyu Shi, Wanqi Shi, Xiaoyu Su, Lianghao Sun, Haozhi Sun, Peiqin Wan, Pengfei Wang, Chao Wang, Chenyu Wang, Meng Wang, Qiulin Wang, Runqi Wang, Xintao Wang, Xuebo Wang, Zekun Wei, Min Wen, Tiancheng Wu, Guohao Wu, Xiaoshi Wu, Zhenhua Xie, Da Xiong, Yingtong Xu, Yulong Yang, Sile Yang, Zikang Ye, Weicai Yuan, Ziyang Zhang, Shenglong Zhang, Shuaiyu Zhang, Yuanxing Zhang, Yufan Zhao, Wenzheng Zhou, Ruiliang Zhou, Yan Zhu, Guosheng Zhu, Yongjie |
| author_facet | Kling Team Chen, Jialu Ci, Yuanzheng Du, Xiangyu Feng, Zipeng Gai, Kun Guo, Sainan Han, Feng He, Jingbin He, Kang Hu, Xiao Hu, Xiaohua Jiang, Boyuan Kong, Fangyuan Li, Hang Li, Jie Li, Qingyu Li, Shen Li, Xiaohan Li, Yan Liang, Jiajun Liao, Borui Liao, Yiqiao Lin, Weihong Liu, Quande Liu, Xiaokun Liu, Yilun Liu, Yuliang Lu, Shun Mao, Hangyu Mao, Yunyao Ouyang, Haodong Qin, Wenyu Shi, Wanqi Shi, Xiaoyu Su, Lianghao Sun, Haozhi Sun, Peiqin Wan, Pengfei Wang, Chao Wang, Chenyu Wang, Meng Wang, Qiulin Wang, Runqi Wang, Xintao Wang, Xuebo Wang, Zekun Wei, Min Wen, Tiancheng Wu, Guohao Wu, Xiaoshi Wu, Zhenhua Xie, Da Xiong, Yingtong Xu, Yulong Yang, Sile Yang, Zikang Ye, Weicai Yuan, Ziyang Zhang, Shenglong Zhang, Shuaiyu Zhang, Yuanxing Zhang, Yufan Zhao, Wenzheng Zhou, Ruiliang Zhou, Yan Zhu, Guosheng Zhu, Yongjie |
| contents | We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional separation among diverse video generation, editing, and intelligent reasoning tasks, integrating them into a holistic system. Unlike disjointed pipeline approaches, Kling-Omni supports a diverse range of user inputs, including text instructions, reference images, and video contexts, processing them into a unified multimodal representation to deliver cinematic-quality and highly-intelligent video content creation. To support these capabilities, we constructed a comprehensive data system that serves as the foundation for multimodal video creation. The framework is further empowered by efficient large-scale pre-training strategies and infrastructure optimizations for inference. Comprehensive evaluations reveal that Kling-Omni demonstrates exceptional capabilities in in-context generation, reasoning-based editing, and multimodal instruction following. Moving beyond a content creation tool, we believe Kling-Omni is a pivotal advancement toward multimodal world simulators capable of perceiving, reasoning, generating and interacting with the dynamic and complex worlds. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_16776 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Kling-Omni Technical Report Kling Team Chen, Jialu Ci, Yuanzheng Du, Xiangyu Feng, Zipeng Gai, Kun Guo, Sainan Han, Feng He, Jingbin He, Kang Hu, Xiao Hu, Xiaohua Jiang, Boyuan Kong, Fangyuan Li, Hang Li, Jie Li, Qingyu Li, Shen Li, Xiaohan Li, Yan Liang, Jiajun Liao, Borui Liao, Yiqiao Lin, Weihong Liu, Quande Liu, Xiaokun Liu, Yilun Liu, Yuliang Lu, Shun Mao, Hangyu Mao, Yunyao Ouyang, Haodong Qin, Wenyu Shi, Wanqi Shi, Xiaoyu Su, Lianghao Sun, Haozhi Sun, Peiqin Wan, Pengfei Wang, Chao Wang, Chenyu Wang, Meng Wang, Qiulin Wang, Runqi Wang, Xintao Wang, Xuebo Wang, Zekun Wei, Min Wen, Tiancheng Wu, Guohao Wu, Xiaoshi Wu, Zhenhua Xie, Da Xiong, Yingtong Xu, Yulong Yang, Sile Yang, Zikang Ye, Weicai Yuan, Ziyang Zhang, Shenglong Zhang, Shuaiyu Zhang, Yuanxing Zhang, Yufan Zhao, Wenzheng Zhou, Ruiliang Zhou, Yan Zhu, Guosheng Zhu, Yongjie Computer Vision and Pattern Recognition We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional separation among diverse video generation, editing, and intelligent reasoning tasks, integrating them into a holistic system. Unlike disjointed pipeline approaches, Kling-Omni supports a diverse range of user inputs, including text instructions, reference images, and video contexts, processing them into a unified multimodal representation to deliver cinematic-quality and highly-intelligent video content creation. To support these capabilities, we constructed a comprehensive data system that serves as the foundation for multimodal video creation. The framework is further empowered by efficient large-scale pre-training strategies and infrastructure optimizations for inference. Comprehensive evaluations reveal that Kling-Omni demonstrates exceptional capabilities in in-context generation, reasoning-based editing, and multimodal instruction following. Moving beyond a content creation tool, we believe Kling-Omni is a pivotal advancement toward multimodal world simulators capable of perceiving, reasoning, generating and interacting with the dynamic and complex worlds. |
| title | Kling-Omni Technical Report |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.16776 |