Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909545236791296 |
|---|---|
| author | Li, Cheng Liu, Jiexiong Chen, Yixuan Jia, Yanqin |
| author_facet | Li, Cheng Liu, Jiexiong Chen, Yixuan Jia, Yanqin |
| contents | In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBaize-VoT-R1 video inference model based on a long-sequence image encoder, along with its training and application methods. By integrating image packing technology, the Autonomy-of-Experts (AoE) architecture, and combining the video of Thought (VoT), a large language model (LLM) trained with large-scale reinforcement learning, and multiple training techniques, the efficiency and accuracy of the model in video inference tasks are effectively improved. Experiments show that this model performs outstandingly in multiple tests, providing a new solution for video-language understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_15807 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture Li, Cheng Liu, Jiexiong Chen, Yixuan Jia, Yanqin Artificial Intelligence In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBaize-VoT-R1 video inference model based on a long-sequence image encoder, along with its training and application methods. By integrating image packing technology, the Autonomy-of-Experts (AoE) architecture, and combining the video of Thought (VoT), a large language model (LLM) trained with large-scale reinforcement learning, and multiple training techniques, the efficiency and accuracy of the model in video inference tasks are effectively improved. Experiments show that this model performs outstandingly in multiple tests, providing a new solution for video-language understanding. |
| title | Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2503.15807 |