Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Jia, Yanqin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909545236791296
author Li, Cheng
Liu, Jiexiong
Chen, Yixuan
Jia, Yanqin
author_facet Li, Cheng
Liu, Jiexiong
Chen, Yixuan
Jia, Yanqin
contents In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBaize-VoT-R1 video inference model based on a long-sequence image encoder, along with its training and application methods. By integrating image packing technology, the Autonomy-of-Experts (AoE) architecture, and combining the video of Thought (VoT), a large language model (LLM) trained with large-scale reinforcement learning, and multiple training techniques, the efficiency and accuracy of the model in video inference tasks are effectively improved. Experiments show that this model performs outstandingly in multiple tests, providing a new solution for video-language understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15807
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
Li, Cheng
Liu, Jiexiong
Chen, Yixuan
Jia, Yanqin
Artificial Intelligence
In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBaize-VoT-R1 video inference model based on a long-sequence image encoder, along with its training and application methods. By integrating image packing technology, the Autonomy-of-Experts (AoE) architecture, and combining the video of Thought (VoT), a large language model (LLM) trained with large-scale reinforcement learning, and multiple training techniques, the efficiency and accuracy of the model in video inference tasks are effectively improved. Experiments show that this model performs outstandingly in multiple tests, providing a new solution for video-language understanding.
title Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
topic Artificial Intelligence
url https://arxiv.org/abs/2503.15807