VQQA: An Agentic Approach for Video Evaluation and Quality Improvement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yiwen, Pfister, Tomas, Song, Yale
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918384624467968
author Song, Yiwen
Pfister, Tomas
Song, Yale
author_facet Song, Yiwen
Pfister, Tomas
Song, Yale
contents Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access to model internals. To address this, we present VQQA (Video Quality Question Answering), a unified, multi-agent framework generalizable across diverse input modalities and video generation tasks. By dynamically generating visual questions and using the resulting Vision-Language Model (VLM) critiques as semantic gradients, VQQA replaces traditional, passive evaluation metrics with human-interpretable, actionable feedback. This enables a highly efficient, closed-loop prompt optimization process via a black-box natural language interface. Extensive experiments demonstrate that VQQA effectively isolates and resolves visual artifacts, substantially improving generation quality in just a few refinement steps. Applicable to both text-to-video (T2V) and image-to-video (I2V) tasks, our method achieves absolute improvements of +11.57% on T2V-CompBench and +8.43% on VBench2 over vanilla generation, significantly outperforming state-of-the-art stochastic search and prompt optimization techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12310
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VQQA: An Agentic Approach for Video Evaluation and Quality Improvement
Song, Yiwen
Pfister, Tomas
Song, Yale
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multiagent Systems
Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access to model internals. To address this, we present VQQA (Video Quality Question Answering), a unified, multi-agent framework generalizable across diverse input modalities and video generation tasks. By dynamically generating visual questions and using the resulting Vision-Language Model (VLM) critiques as semantic gradients, VQQA replaces traditional, passive evaluation metrics with human-interpretable, actionable feedback. This enables a highly efficient, closed-loop prompt optimization process via a black-box natural language interface. Extensive experiments demonstrate that VQQA effectively isolates and resolves visual artifacts, substantially improving generation quality in just a few refinement steps. Applicable to both text-to-video (T2V) and image-to-video (I2V) tasks, our method achieves absolute improvements of +11.57% on T2V-CompBench and +8.43% on VBench2 over vanilla generation, significantly outperforming state-of-the-art stochastic search and prompt optimization techniques.
title VQQA: An Agentic Approach for Video Evaluation and Quality Improvement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2603.12310