Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haoji, Gu, Xin, Li, Jiawen, Ma, Chixiang, Bai, Sule, Zhang, Chubin, Zhang, Bowen, Zhou, Zhichao, He, Dongliang, Tang, Yansong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908516227219456
author Zhang, Haoji
Gu, Xin
Li, Jiawen
Ma, Chixiang
Bai, Sule
Zhang, Chubin
Zhang, Bowen
Zhou, Zhichao
He, Dongliang
Tang, Yansong
author_facet Zhang, Haoji
Gu, Xin
Li, Jiawen
Ma, Chixiang
Bai, Sule
Zhang, Chubin
Zhang, Bowen
Zhou, Zhichao
He, Dongliang
Tang, Yansong
contents The video reasoning ability of multimodal large language models (MLLMs) is crucial for downstream tasks like video question answering and temporal grounding. While recent approaches have explored text-based chain-of-thought (CoT) reasoning for MLLMs, these methods often suffer from limited cross-modal interaction and increased hallucination, especially with longer videos or reasoning chains. To address these challenges, we propose Video Intelligence via Tool-Augmented Learning (VITAL), a novel end-to-end agentic video reasoning framework. With a visual toolbox, the model can densely sample new video frames on demand and generate multimodal CoT for precise long video reasoning. We observe that temporal grounding and question answering are mutually beneficial for video understanding tasks. Therefore, we construct two high-quality multi-task video reasoning datasets MTVR-CoT-72k for supervised fine-tuning and MTVR-RL-110k for reinforcement learning. Moreover, we propose a Difficulty-aware Group Relative Policy Optimization algorithm (DGRPO) to mitigate difficulty imbalance in multi-task reinforcement learning. Extensive experiments on 11 challenging video understanding benchmarks demonstrate the advanced reasoning ability of VITAL, outperforming existing methods in video question answering and temporal grounding tasks, especially in long video scenarios. Code is available at https://zhang9302002.github.io/thinkingwithvideos-page/.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04416
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Zhang, Haoji
Gu, Xin
Li, Jiawen
Ma, Chixiang
Bai, Sule
Zhang, Chubin
Zhang, Bowen
Zhou, Zhichao
He, Dongliang
Tang, Yansong
Computer Vision and Pattern Recognition
The video reasoning ability of multimodal large language models (MLLMs) is crucial for downstream tasks like video question answering and temporal grounding. While recent approaches have explored text-based chain-of-thought (CoT) reasoning for MLLMs, these methods often suffer from limited cross-modal interaction and increased hallucination, especially with longer videos or reasoning chains. To address these challenges, we propose Video Intelligence via Tool-Augmented Learning (VITAL), a novel end-to-end agentic video reasoning framework. With a visual toolbox, the model can densely sample new video frames on demand and generate multimodal CoT for precise long video reasoning. We observe that temporal grounding and question answering are mutually beneficial for video understanding tasks. Therefore, we construct two high-quality multi-task video reasoning datasets MTVR-CoT-72k for supervised fine-tuning and MTVR-RL-110k for reinforcement learning. Moreover, we propose a Difficulty-aware Group Relative Policy Optimization algorithm (DGRPO) to mitigate difficulty imbalance in multi-task reinforcement learning. Extensive experiments on 11 challenging video understanding benchmarks demonstrate the advanced reasoning ability of VITAL, outperforming existing methods in video question answering and temporal grounding tasks, especially in long video scenarios. Code is available at https://zhang9302002.github.io/thinkingwithvideos-page/.
title Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.04416