Understanding Complexity in VideoQA via Visual Program Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Eyzaguirre, Cristobal, Vasiljevic, Igor, Dave, Achal, Wu, Jiajun, Ambrus, Rares Andrei, Kollar, Thomas, Niebles, Juan Carlos, Tokmakov, Pavel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908370750930944
author Eyzaguirre, Cristobal
Vasiljevic, Igor
Dave, Achal
Wu, Jiajun
Ambrus, Rares Andrei
Kollar, Thomas
Niebles, Juan Carlos
Tokmakov, Pavel
author_facet Eyzaguirre, Cristobal
Vasiljevic, Igor
Dave, Achal
Wu, Jiajun
Ambrus, Rares Andrei
Kollar, Thomas
Niebles, Juan Carlos
Tokmakov, Pavel
contents We propose a data-driven approach to analyzing query complexity in Video Question Answering (VideoQA). Previous efforts in benchmark design have relied on human expertise to design challenging questions, yet we experimentally show that humans struggle to predict which questions are difficult for machine learning models. Our automatic approach leverages recent advances in code generation for visual question answering, using the complexity of generated code as a proxy for question difficulty. We demonstrate that this measure correlates significantly better with model performance than human estimates. To operationalize this insight, we propose an algorithm for estimating question complexity from code. It identifies fine-grained primitives that correlate with the hardest questions for any given set of models, making it easy to scale to new approaches in the future. Finally, to further illustrate the utility of our method, we extend it to automatically generate complex questions, constructing a new benchmark that is 1.9 times harder than the popular NExT-QA.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13429
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Complexity in VideoQA via Visual Program Generation
Eyzaguirre, Cristobal
Vasiljevic, Igor
Dave, Achal
Wu, Jiajun
Ambrus, Rares Andrei
Kollar, Thomas
Niebles, Juan Carlos
Tokmakov, Pavel
Computer Vision and Pattern Recognition
We propose a data-driven approach to analyzing query complexity in Video Question Answering (VideoQA). Previous efforts in benchmark design have relied on human expertise to design challenging questions, yet we experimentally show that humans struggle to predict which questions are difficult for machine learning models. Our automatic approach leverages recent advances in code generation for visual question answering, using the complexity of generated code as a proxy for question difficulty. We demonstrate that this measure correlates significantly better with model performance than human estimates. To operationalize this insight, we propose an algorithm for estimating question complexity from code. It identifies fine-grained primitives that correlate with the hardest questions for any given set of models, making it easy to scale to new approaches in the future. Finally, to further illustrate the utility of our method, we extend it to automatically generate complex questions, constructing a new benchmark that is 1.9 times harder than the popular NExT-QA.
title Understanding Complexity in VideoQA via Visual Program Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.13429