ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiao, Liu, Nayu, Zhu, Junnan, Chen, Ruirui, Xiang, Guohui, Wang, Changjian, Wei, Kaiwen, Li, Rongzhen, Zhong, Jiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
by: Feng, Jiazhan, et al.
Published: (2025)
by: Feng, Jiazhan, et al.
Published: (2025)
Aurora: Unified Video Editing with a Tool-Using Agent
by: Yu, Yongsheng, et al.
Published: (2026)
by: Yu, Yongsheng, et al.
Published: (2026)
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026)
by: Lin, Jingyang, et al.
Published: (2026)
AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmark
by: Xiang, Xinhao, et al.
Published: (2025)
by: Xiang, Xinhao, et al.
Published: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
by: Zhang, Haoji, et al.
Published: (2025)
by: Zhang, Haoji, et al.
Published: (2025)
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
by: Ye, Zi, et al.
Published: (2026)
by: Ye, Zi, et al.
Published: (2026)
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
by: Shabbir, Akashah, et al.
Published: (2026)
by: Shabbir, Akashah, et al.
Published: (2026)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
WALT: Web Agents that Learn Tools
by: Prabhu, Viraj, et al.
Published: (2025)
by: Prabhu, Viraj, et al.
Published: (2025)
OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation
by: Zhao, Sijie, et al.
Published: (2026)
by: Zhao, Sijie, et al.
Published: (2026)
Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
by: Guo, Ruocheng, et al.
Published: (2026)
by: Guo, Ruocheng, et al.
Published: (2026)
TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents
by: Yu, Bihui, et al.
Published: (2026)
by: Yu, Bihui, et al.
Published: (2026)
COLT: Enhancing Video Large Language Models with Continual Tool Usage
by: Liu, Yuyang, et al.
Published: (2025)
by: Liu, Yuyang, et al.
Published: (2025)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
by: Gao, Zhe, et al.
Published: (2026)
by: Gao, Zhe, et al.
Published: (2026)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
by: Wang, Yan, et al.
Published: (2024)
by: Wang, Yan, et al.
Published: (2024)
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
by: Yuan, Zhenlong, et al.
Published: (2025)
by: Yuan, Zhenlong, et al.
Published: (2025)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
by: Tian, Shulin, et al.
Published: (2025)
by: Tian, Shulin, et al.
Published: (2025)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
by: Shabbir, Akashah, et al.
Published: (2025)
by: Shabbir, Akashah, et al.
Published: (2025)
PoseTraj: Pose-Aware Trajectory Control in Video Diffusion
by: Ji, Longbin, et al.
Published: (2025)
by: Ji, Longbin, et al.
Published: (2025)
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation
by: Jiang, Xi, et al.
Published: (2026)
by: Jiang, Xi, et al.
Published: (2026)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
by: Jia, Hongrui, et al.
Published: (2025)
by: Jia, Hongrui, et al.
Published: (2025)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
by: Zhong, Yaoyao, et al.
Published: (2023)
by: Zhong, Yaoyao, et al.
Published: (2023)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026)
by: Zhao, Yiming, et al.
Published: (2026)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
by: Wang, Boran, et al.
Published: (2025)
by: Wang, Boran, et al.
Published: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
by: Zhang, Xiaoyi, et al.
Published: (2025)
by: Zhang, Xiaoyi, et al.
Published: (2025)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
by: Hannan, Tanveer, et al.
Published: (2024)
by: Hannan, Tanveer, et al.
Published: (2024)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
by: Qian, Kangan, et al.
Published: (2025)
by: Qian, Kangan, et al.
Published: (2025)
GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Similar Items
-
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
by: Feng, Jiazhan, et al.
Published: (2025) -
Aurora: Unified Video Editing with a Tool-Using Agent
by: Yu, Yongsheng, et al.
Published: (2026) -
VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
by: Lin, Jingyang, et al.
Published: (2026) -
AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmark
by: Xiang, Xinhao, et al.
Published: (2025) -
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
by: Zhang, Haoji, et al.
Published: (2025)