Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Chengwen, Yu, Xiaomin, Chang, Zhuoyue, Huang, Zhe, Zhang, Shuo, Lian, Heng, Dang, Jisheng, Xu, Rui, Hu, Sen, Hou, Jianheng, Qin, Chengwei, Hu, Xiaobin, Wang, Kunyi, Yang, Zhi, Peng, Hao, Peng, Hong, Chen, Ronghao, Wang, Huacan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916022094659584
author Liu, Chengwen
Yu, Xiaomin
Chang, Zhuoyue
Huang, Zhe
Zhang, Shuo
Lian, Heng
Dang, Jisheng
Xu, Rui
Hu, Sen
Hou, Jianheng
Qin, Chengwei
Hu, Xiaobin
Wang, Kunyi
Yang, Zhi
Peng, Hao
Peng, Hong
Chen, Ronghao
Wang, Huacan
author_facet Liu, Chengwen
Yu, Xiaomin
Chang, Zhuoyue
Huang, Zhe
Zhang, Shuo
Lian, Heng
Dang, Jisheng
Xu, Rui
Hu, Sen
Hou, Jianheng
Qin, Chengwei
Hu, Xiaobin
Wang, Kunyi
Yang, Zhi
Peng, Hao
Peng, Hong
Chen, Ronghao
Wang, Huacan
contents In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore need to jointly perform cross-frame clue extraction, iterative retrieval, and multi-hop reasoning-based verification. To bridge this gap, we construct the first video deep research benchmark, VideoDR. VideoDR centers on video-conditioned open-domain video question answering, requiring cross-frame visual anchor extraction, interactive web retrieval, and multi-hop reasoning over joint video-web evidence; through rigorous human annotation and quality control, we obtain high-quality video deep research samples spanning six semantic domains. We evaluate multiple closed-source and open-source multimodal large language models under both the Workflow and Agentic paradigms, and the results show that Agentic is not consistently superior to Workflow: its gains depend on a model's ability to maintain the initial video anchors over long retrieval chains. Further analysis indicates that goal drift and long-horizon consistency are the core bottlenecks. In sum, VideoDR provides a systematic benchmark for studying video agents in open-web settings and reveals the key challenges for next-generation video deep research agents.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06943
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
Liu, Chengwen
Yu, Xiaomin
Chang, Zhuoyue
Huang, Zhe
Zhang, Shuo
Lian, Heng
Dang, Jisheng
Xu, Rui
Hu, Sen
Hou, Jianheng
Qin, Chengwei
Hu, Xiaobin
Wang, Kunyi
Yang, Zhi
Peng, Hao
Peng, Hong
Chen, Ronghao
Wang, Huacan
Computer Vision and Pattern Recognition
Artificial Intelligence
In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed across the open web; models therefore need to jointly perform cross-frame clue extraction, iterative retrieval, and multi-hop reasoning-based verification. To bridge this gap, we construct the first video deep research benchmark, VideoDR. VideoDR centers on video-conditioned open-domain video question answering, requiring cross-frame visual anchor extraction, interactive web retrieval, and multi-hop reasoning over joint video-web evidence; through rigorous human annotation and quality control, we obtain high-quality video deep research samples spanning six semantic domains. We evaluate multiple closed-source and open-source multimodal large language models under both the Workflow and Agentic paradigms, and the results show that Agentic is not consistently superior to Workflow: its gains depend on a model's ability to maintain the initial video anchors over long retrieval chains. Further analysis indicates that goal drift and long-horizon consistency are the core bottlenecks. In sum, VideoDR provides a systematic benchmark for studying video agents in open-web settings and reveals the key challenges for next-generation video deep research agents.
title Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.06943