Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Chengwen, Yu, Xiaomin, Chang, Zhuoyue, Huang, Zhe, Zhang, Shuo, Lian, Heng, Dang, Jisheng, Xu, Rui, Hu, Sen, Hou, Jianheng, Qin, Chengwei, Hu, Xiaobin, Wang, Kunyi, Yang, Zhi, Peng, Hao, Peng, Hong, Chen, Ronghao, Wang, Huacan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
by: Wang, Qihao, et al.
Published: (2026)
by: Wang, Qihao, et al.
Published: (2026)
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
by: Zhang, Congzhi, et al.
Published: (2025)
by: Zhang, Congzhi, et al.
Published: (2025)
Reinforcing Video Reasoning with Focused Thinking
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
by: Jisheng, Dang, et al.
Published: (2025)
by: Jisheng, Dang, et al.
Published: (2025)
QuantaAlpha: An Evolutionary Framework for LLM-Driven Alpha Mining
by: Han, Jun, et al.
Published: (2026)
by: Han, Jun, et al.
Published: (2026)
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
by: Jiang, Tianyi, et al.
Published: (2026)
by: Jiang, Tianyi, et al.
Published: (2026)
CloneMem: Benchmarking Long-Term Memory for AI Clones
by: Hu, Sen, et al.
Published: (2026)
by: Hu, Sen, et al.
Published: (2026)
SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
by: Lin, Jiaye, et al.
Published: (2025)
by: Lin, Jiaye, et al.
Published: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
by: Zhang, Weizhi, et al.
Published: (2025)
by: Zhang, Weizhi, et al.
Published: (2025)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
by: Li, Kuan, et al.
Published: (2026)
by: Li, Kuan, et al.
Published: (2026)
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
by: Wang, Wentao, et al.
Published: (2025)
by: Wang, Wentao, et al.
Published: (2025)
VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
EpochX: Building the Infrastructure for an Emergent Agent Civilization
by: Wang, Huacan, et al.
Published: (2026)
by: Wang, Huacan, et al.
Published: (2026)
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
by: He, Qingdong, et al.
Published: (2026)
by: He, Qingdong, et al.
Published: (2026)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification
by: Hu, Qisheng, et al.
Published: (2025)
by: Hu, Qisheng, et al.
Published: (2025)
Does Memory Need Graphs? A Unified Framework and Empirical Analysis for Long-Term Dialog Memory
by: Hu, Sen, et al.
Published: (2026)
by: Hu, Sen, et al.
Published: (2026)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
by: Zhu, Minghao, et al.
Published: (2024)
by: Zhu, Minghao, et al.
Published: (2024)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
by: Wu, Tingyu, et al.
Published: (2026)
by: Wu, Tingyu, et al.
Published: (2026)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
by: Hu, X., et al.
Published: (2025)
by: Hu, X., et al.
Published: (2025)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
by: Yin, Xinlei, et al.
Published: (2026)
by: Yin, Xinlei, et al.
Published: (2026)
Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
by: Dai, Chengwei, et al.
Published: (2024)
by: Dai, Chengwei, et al.
Published: (2024)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
Explicit Uncertainty Modeling for Video Watch Time Prediction
by: Wu, Shanshan, et al.
Published: (2025)
by: Wu, Shanshan, et al.
Published: (2025)
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)
by: Peng, Wenshuo, et al.
Published: (2025)
End-to-end Semantic-centric Video-based Multimodal Affective Computing
by: Lin, Ronghao, et al.
Published: (2024)
by: Lin, Ronghao, et al.
Published: (2024)
Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding
by: Zheng, Muyang, et al.
Published: (2026)
by: Zheng, Muyang, et al.
Published: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
by: Huang, Xijie, et al.
Published: (2026)
by: Huang, Xijie, et al.
Published: (2026)
Similar Items
-
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025) -
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
by: Dang, Jisheng, et al.
Published: (2025) -
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
by: Wang, Qihao, et al.
Published: (2026) -
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
by: Zhang, Congzhi, et al.
Published: (2025) -
Reinforcing Video Reasoning with Focused Thinking
by: Dang, Jisheng, et al.
Published: (2025)