AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patel, Urjitkumar, Yeh, Fang-Chun, Gondhalekar, Chinmay
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914462286479360
author Patel, Urjitkumar
Yeh, Fang-Chun
Gondhalekar, Chinmay
author_facet Patel, Urjitkumar
Yeh, Fang-Chun
Gondhalekar, Chinmay
contents With the increasing prevalence of video content, effectively understanding and answering questions about long form videos has become essential for numerous applications. Although large vision language models (LVLMs) have enhanced performance, they often face challenges with nuanced queries that demand both a comprehensive understanding and detailed analysis. To overcome these obstacles, we introduce AVATAAR, a modular and interpretable framework that combines global and local video context, along with a Pre Retrieval Thinking Agent and a Rethink Module. AVATAAR creates a persistent global summary and establishes a feedback loop between the Rethink Module and the Pre Retrieval Thinking Agent, allowing the system to refine its retrieval strategies based on partial answers and replicate human-like iterative reasoning. On the CinePile benchmark, AVATAAR demonstrates significant improvements over a baseline, achieving relative gains of +5.6% in temporal reasoning, +5% in technical queries, +8% in theme-based questions, and +8.2% in narrative comprehension. Our experiments confirm that each module contributes positively to the overall performance, with the feedback loop being crucial for adaptability. These findings highlight AVATAAR's effectiveness in enhancing video understanding capabilities. Ultimately, AVATAAR presents a scalable solution for long-form Video Question Answering (QA), merging accuracy, interpretability, and extensibility.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
Patel, Urjitkumar
Yeh, Fang-Chun
Gondhalekar, Chinmay
Computer Vision and Pattern Recognition
68T45, 68T50 (Primary) 68T10, 68T07, 62H35 (Secondary)
I.2.10; I.2.7; H.3.3; I.4.8; I.5.4
With the increasing prevalence of video content, effectively understanding and answering questions about long form videos has become essential for numerous applications. Although large vision language models (LVLMs) have enhanced performance, they often face challenges with nuanced queries that demand both a comprehensive understanding and detailed analysis. To overcome these obstacles, we introduce AVATAAR, a modular and interpretable framework that combines global and local video context, along with a Pre Retrieval Thinking Agent and a Rethink Module. AVATAAR creates a persistent global summary and establishes a feedback loop between the Rethink Module and the Pre Retrieval Thinking Agent, allowing the system to refine its retrieval strategies based on partial answers and replicate human-like iterative reasoning. On the CinePile benchmark, AVATAAR demonstrates significant improvements over a baseline, achieving relative gains of +5.6% in temporal reasoning, +5% in technical queries, +8% in theme-based questions, and +8.2% in narrative comprehension. Our experiments confirm that each module contributes positively to the overall performance, with the feedback loop being crucial for adaptability. These findings highlight AVATAAR's effectiveness in enhancing video understanding capabilities. Ultimately, AVATAAR presents a scalable solution for long-form Video Question Answering (QA), merging accuracy, interpretability, and extensibility.
title AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
topic Computer Vision and Pattern Recognition
68T45, 68T50 (Primary) 68T10, 68T07, 62H35 (Secondary)
I.2.10; I.2.7; H.3.3; I.4.8; I.5.4
url https://arxiv.org/abs/2511.15578