Understanding Complexity in VideoQA via Visual Program Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Eyzaguirre, Cristobal, Vasiljevic, Igor, Dave, Achal, Wu, Jiajun, Ambrus, Rares Andrei, Kollar, Thomas, Niebles, Juan Carlos, Tokmakov, Pavel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
by: Guizilini, Vitor, et al.
Published: (2024)
by: Guizilini, Vitor, et al.
Published: (2024)
Understanding Video Transformers via Universal Concept Discovery
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model
by: Yu, Keunwoo Peter, et al.
Published: (2024)
by: Yu, Keunwoo Peter, et al.
Published: (2024)
Video Generators are Robot Policies
by: Liang, Junbang, et al.
Published: (2025)
by: Liang, Junbang, et al.
Published: (2025)
Linearizing Large Language Models
by: Mercat, Jean, et al.
Published: (2024)
by: Mercat, Jean, et al.
Published: (2024)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
by: Liang, Junbang, et al.
Published: (2024)
by: Liang, Junbang, et al.
Published: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
Streaming Detection of Queried Event Start
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
by: Eyzaguirre, Cristobal, et al.
Published: (2024)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
by: Guo, Jiangyuan, et al.
Published: (2024)
by: Guo, Jiangyuan, et al.
Published: (2024)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
by: Liu, Yunong, et al.
Published: (2024)
by: Liu, Yunong, et al.
Published: (2024)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
by: Song, Zijie, et al.
Published: (2025)
by: Song, Zijie, et al.
Published: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
by: He, Zhixian, et al.
Published: (2024)
by: He, Zhixian, et al.
Published: (2024)
Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models
by: Chu, Wen-Hsuan, et al.
Published: (2023)
by: Chu, Wen-Hsuan, et al.
Published: (2023)
pix2gestalt: Amodal Segmentation by Synthesizing Wholes
by: Ozguroglu, Ege, et al.
Published: (2024)
by: Ozguroglu, Ege, et al.
Published: (2024)
Incorporating dense metric depth into neural 3D representations for view synthesis and relighting
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
by: Chaudhury, Arkadeep Narayan, et al.
Published: (2024)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025)
by: Parikh, Chirag, et al.
Published: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Should VLMs be Pre-trained with Image Data?
by: Keh, Sedrick, et al.
Published: (2025)
by: Keh, Sedrick, et al.
Published: (2025)
CARTO: Category and Joint Agnostic Reconstruction of ARTiculated Objects
by: Heppert, Nick, et al.
Published: (2023)
by: Heppert, Nick, et al.
Published: (2023)
AllTracker: Efficient Dense Point Tracking at High Resolution
by: Harley, Adam W., et al.
Published: (2025)
by: Harley, Adam W., et al.
Published: (2025)
EXPANSIÓN Y LÍMITES DE LA BUENA FE OBJETIVA – A PROPÓSITO DEL “PROYECTO DE PRINCIPIOS LATINOAMERICANOS DE DERECHO DE LOS CONTRATOS”
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
by: Cristóbal Eyzaguirre Baeza
Published: (2013)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
by: Van Hoorick, Basile, et al.
Published: (2024)
by: Van Hoorick, Basile, et al.
Published: (2024)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
by: Heyward, Joseph, et al.
Published: (2024)
by: Heyward, Joseph, et al.
Published: (2024)
Transcrib3D: 3D Referring Expression Resolution through Large Language Models
by: Fang, Jiading, et al.
Published: (2024)
by: Fang, Jiading, et al.
Published: (2024)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
by: Hu, Yuhang, et al.
Published: (2025)
by: Hu, Yuhang, et al.
Published: (2025)
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
ReFiNe: Recursive Field Networks for Cross-modal Multi-scene Representation
by: Zakharov, Sergey, et al.
Published: (2024)
by: Zakharov, Sergey, et al.
Published: (2024)
Capturing Visual Environment Structure Correlates with Control Performance
by: Dong, Jiahua, et al.
Published: (2026)
by: Dong, Jiahua, et al.
Published: (2026)
Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry
by: Kanai, Takayuki, et al.
Published: (2024)
by: Kanai, Takayuki, et al.
Published: (2024)
Similar Items
-
GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion
by: Guizilini, Vitor, et al.
Published: (2024) -
Understanding Video Transformers via Universal Concept Discovery
by: Kowal, Matthew, et al.
Published: (2024) -
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026) -
Espresso: High Compression For Rich Extraction From Videos for Your Vision-Language Model
by: Yu, Keunwoo Peter, et al.
Published: (2024) -
Video Generators are Robot Policies
by: Liang, Junbang, et al.
Published: (2025)