VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chenglin, Chen, Qianglong, Han, Feng, Wang, Yikun, Yin, Xingxi, Gong, Yan, Li, Ruilin, Zhang, Yin, Wang, Jiaqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908977401430016
author Li, Chenglin
Chen, Qianglong
Han, Feng
Wang, Yikun
Yin, Xingxi
Gong, Yan
Li, Ruilin
Zhang, Yin
Wang, Jiaqi
author_facet Li, Chenglin
Chen, Qianglong
Han, Feng
Wang, Yikun
Yin, Xingxi
Gong, Yan
Li, Ruilin
Zhang, Yin
Wang, Jiaqi
contents Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames, which weakens temporal localization and leads to substantial information loss in long videos. Agentic tools such as temporal retrieval, spatial zoom, and temporal zoom offer a natural way to overcome these limitations by enabling adaptive exploration of key moments. However, constructing agentic video understanding data requires models that already possess strong long-form video comprehension, creating a circular dependency. We address this challenge with VideoThinker, an agentic Video Large Language Model trained entirely on synthetic tool interaction trajectories. Our key idea is to convert videos into rich captions and employ a powerful agentic language model to generate multi-step tool use sequences in caption space. These trajectories are subsequently grounded back to video by replacing captions with the corresponding frames, yielding a large-scale interleaved video and tool reasoning dataset without requiring any long-form understanding from the underlying model. Training on this synthetic agentic dataset equips VideoThinker with dynamic reasoning capabilities, adaptive temporal exploration, and multi-step tool use. Remarkably, VideoThinker significantly outperforms both caption-only language model agents and strong video model baselines across long-video benchmarks, demonstrating the effectiveness of tool augmented synthetic data and adaptive retrieval and zoom reasoning for long-form video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15724
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
Li, Chenglin
Chen, Qianglong
Han, Feng
Wang, Yikun
Yin, Xingxi
Gong, Yan
Li, Ruilin
Zhang, Yin
Wang, Jiaqi
Computer Vision and Pattern Recognition
Artificial Intelligence
Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames, which weakens temporal localization and leads to substantial information loss in long videos. Agentic tools such as temporal retrieval, spatial zoom, and temporal zoom offer a natural way to overcome these limitations by enabling adaptive exploration of key moments. However, constructing agentic video understanding data requires models that already possess strong long-form video comprehension, creating a circular dependency. We address this challenge with VideoThinker, an agentic Video Large Language Model trained entirely on synthetic tool interaction trajectories. Our key idea is to convert videos into rich captions and employ a powerful agentic language model to generate multi-step tool use sequences in caption space. These trajectories are subsequently grounded back to video by replacing captions with the corresponding frames, yielding a large-scale interleaved video and tool reasoning dataset without requiring any long-form understanding from the underlying model. Training on this synthetic agentic dataset equips VideoThinker with dynamic reasoning capabilities, adaptive temporal exploration, and multi-step tool use. Remarkably, VideoThinker significantly outperforms both caption-only language model agents and strong video model baselines across long-video benchmarks, demonstrating the effectiveness of tool augmented synthetic data and adaptive retrieval and zoom reasoning for long-form video understanding.
title VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.15724