VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aparcedo, Alejandro, Kumar, Akash, Garg, Aaryan, Pham, Dalton, Chen, Wen-Kai, Bharadwaj, Anirudh, Chadha, Aman, Rawat, Yogesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913083544305664
author Aparcedo, Alejandro
Kumar, Akash
Garg, Aaryan
Pham, Dalton
Chen, Wen-Kai
Bharadwaj, Anirudh
Chadha, Aman
Rawat, Yogesh
author_facet Aparcedo, Alejandro
Kumar, Akash
Garg, Aaryan
Pham, Dalton
Chen, Wen-Kai
Bharadwaj, Anirudh
Chadha, Aman
Rawat, Yogesh
contents Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action interactions between diverse entities which characterize real-world video understanding. Furthermore, the lack of a systematic framework for analyzing model failures across complementary spatio-temporal axes hinders comprehensive evaluation. To address these gaps, we introduce VISTA, a Video Interaction Spatio-Temporal Analysis benchmark designed for open-set, multi-entity and multi-action spatio-temporal understanding in VLMs. VISTA decomposes videos into interpretable entities, their associated actions, and relational dynamics, enabling multi-axis diagnostics and unified assessment of relational, spatial, and temporal understanding. Our benchmark integrates multiple datasets into a single interaction-aware taxonomy and comprises ~12K curated video-query pairs spanning diverse scenes and complexities. We systematically evaluate 11 state-of-the-art VLMs on VISTA, and break down aggregate performance across our taxonomy to reveal shortcomings and pronounced spatio-temporal biases obscured by traditional metrics. By providing detailed, taxonomy-driven diagnostics on a challenging dataset, VISTA offers a nuanced framework to guide advances in model design, pretraining strategies, and evaluation protocols. Overall, VISTA is the first, large-scale, interaction-aware diagnostic benchmark for spatio-temporal understanding in VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01391
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
Aparcedo, Alejandro
Kumar, Akash
Garg, Aaryan
Pham, Dalton
Chen, Wen-Kai
Bharadwaj, Anirudh
Chadha, Aman
Rawat, Yogesh
Computer Vision and Pattern Recognition
Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action interactions between diverse entities which characterize real-world video understanding. Furthermore, the lack of a systematic framework for analyzing model failures across complementary spatio-temporal axes hinders comprehensive evaluation. To address these gaps, we introduce VISTA, a Video Interaction Spatio-Temporal Analysis benchmark designed for open-set, multi-entity and multi-action spatio-temporal understanding in VLMs. VISTA decomposes videos into interpretable entities, their associated actions, and relational dynamics, enabling multi-axis diagnostics and unified assessment of relational, spatial, and temporal understanding. Our benchmark integrates multiple datasets into a single interaction-aware taxonomy and comprises ~12K curated video-query pairs spanning diverse scenes and complexities. We systematically evaluate 11 state-of-the-art VLMs on VISTA, and break down aggregate performance across our taxonomy to reveal shortcomings and pronounced spatio-temporal biases obscured by traditional metrics. By providing detailed, taxonomy-driven diagnostics on a challenging dataset, VISTA offers a nuanced framework to guide advances in model design, pretraining strategies, and evaluation protocols. Overall, VISTA is the first, large-scale, interaction-aware diagnostic benchmark for spatio-temporal understanding in VLMs.
title VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.01391