See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Zixuan, Peng, Baoyun, Wang, Yufei, Liu, Lin, Dong, Xinxin, Cao, Yunlong, Wang, Xiaodong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915461926486016
author Dong, Zixuan
Peng, Baoyun
Wang, Yufei
Liu, Lin
Dong, Xinxin
Cao, Yunlong
Wang, Xiaodong
author_facet Dong, Zixuan
Peng, Baoyun
Wang, Yufei
Liu, Lin
Dong, Xinxin
Cao, Yunlong
Wang, Xiaodong
contents Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video question answering systems employ rigid pipelines that decouple reasoning from perception, leading to either information loss through premature visual abstraction or computational inefficiency through exhaustive processing. The core limitation lies in the inability to adapt visual extraction to specific reasoning requirements, different queries demand fundamentally different visual evidence from the same video content. In this work, we present CAVIA, a training-free framework that revolutionizes video understanding through reasoning, perception coordination. Unlike conventional approaches where visual processing operates independently of reasoning, CAVIA creates a closed-loop system where reasoning continuously guides visual extraction based on identified information gaps. CAVIA introduces three innovations: (1) hierarchical reasoning, guided localization to precise frames; (2) cross-modal semantic bridging for targeted extraction; (3) confidence-driven iterative synthesis. CAVIA achieves state-of-the-art performance on challenging benchmarks: EgoSchema (65.7%, +5.3%), NExT-QA (76.1%, +2.6%), and IntentQA (73.8%, +6.9%), demonstrating that dynamic reasoning-perception coordination provides a scalable paradigm for video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2508_17932
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
Dong, Zixuan
Peng, Baoyun
Wang, Yufei
Liu, Lin
Dong, Xinxin
Cao, Yunlong
Wang, Xiaodong
Computer Vision and Pattern Recognition
Artificial Intelligence
68T45, 68T05
H.5.1; I.2.10; I.4.8; I.5.4
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video question answering systems employ rigid pipelines that decouple reasoning from perception, leading to either information loss through premature visual abstraction or computational inefficiency through exhaustive processing. The core limitation lies in the inability to adapt visual extraction to specific reasoning requirements, different queries demand fundamentally different visual evidence from the same video content. In this work, we present CAVIA, a training-free framework that revolutionizes video understanding through reasoning, perception coordination. Unlike conventional approaches where visual processing operates independently of reasoning, CAVIA creates a closed-loop system where reasoning continuously guides visual extraction based on identified information gaps. CAVIA introduces three innovations: (1) hierarchical reasoning, guided localization to precise frames; (2) cross-modal semantic bridging for targeted extraction; (3) confidence-driven iterative synthesis. CAVIA achieves state-of-the-art performance on challenging benchmarks: EgoSchema (65.7%, +5.3%), NExT-QA (76.1%, +2.6%), and IntentQA (73.8%, +6.9%), demonstrating that dynamic reasoning-perception coordination provides a scalable paradigm for video understanding.
title See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
topic Computer Vision and Pattern Recognition
Artificial Intelligence
68T45, 68T05
H.5.1; I.2.10; I.4.8; I.5.4
url https://arxiv.org/abs/2508.17932