Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
Fuente:
arXiv
Saved in:
| Main Authors: | Benschop, Pascal, Dauwels, Justin, van Gemert, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of Vision-LLMs in Surveillance Video
by: Benschop, Pascal, et al.
Published: (2025)
by: Benschop, Pascal, et al.
Published: (2025)
Identifying Ethical Biases in Action Recognition Models
by: Baltaretu, Ana, et al.
Published: (2026)
by: Baltaretu, Ana, et al.
Published: (2026)
HAVANA: Hierarchical stochastic neighbor embedding for Accelerated Video ANnotAtions
by: Bobe, Alexandru, et al.
Published: (2024)
by: Bobe, Alexandru, et al.
Published: (2024)
Compositional Scene Understanding through Inverse Generative Modeling
by: Wang, Yanbo, et al.
Published: (2025)
by: Wang, Yanbo, et al.
Published: (2025)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)
by: Wang, Yanbo, et al.
Published: (2023)
Rare-Aware Autoencoding: Reconstructing Spatially Imbalanced Data
by: Garcia, Alejandro Castañeda, et al.
Published: (2026)
by: Garcia, Alejandro Castañeda, et al.
Published: (2026)
End-to-End Chess Recognition
by: Masouris, Athanasios, et al.
Published: (2023)
by: Masouris, Athanasios, et al.
Published: (2023)
End-to-End Implicit Neural Representations for Classification
by: Gielisse, Alexander, et al.
Published: (2025)
by: Gielisse, Alexander, et al.
Published: (2025)
Learning to Adapt to Position Bias in Vision Transformer Classifiers
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
by: Bruintjes, Robert-Jan, et al.
Published: (2025)
LayoutGKN: Graph Similarity Learning of Floor Plans
by: van Engelenburg, Casper, et al.
Published: (2025)
by: van Engelenburg, Casper, et al.
Published: (2025)
Local Attention Transformers for High-Detail Optical Flow Upsampling
by: Gielisse, Alexander, et al.
Published: (2024)
by: Gielisse, Alexander, et al.
Published: (2024)
Pushing the boundaries of event subsampling in event-based video classification using CNNs
by: Araghi, Hesam, et al.
Published: (2024)
by: Araghi, Hesam, et al.
Published: (2024)
Making Every Event Count: Balancing Data Efficiency and Accuracy in Event Camera Subsampling
by: Araghi, Hesam, et al.
Published: (2025)
by: Araghi, Hesam, et al.
Published: (2025)
Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical Systems
by: Garcia, Alejandro Castañeda, et al.
Published: (2024)
by: Garcia, Alejandro Castañeda, et al.
Published: (2024)
Bringing a Personal Point of View: Evaluating Dynamic 3D Gaussian Splatting for Egocentric Scene Reconstruction
by: Warchocki, Jan, et al.
Published: (2026)
by: Warchocki, Jan, et al.
Published: (2026)
Deep Continuous Networks
by: Tomen, Nergis, et al.
Published: (2024)
by: Tomen, Nergis, et al.
Published: (2024)
ARC: Anchored Representation Clouds for High-Resolution INR Classification
by: Luijmes, Joost, et al.
Published: (2025)
by: Luijmes, Joost, et al.
Published: (2025)
Deep activity propagation via weight initialization in spiking neural networks
by: Micheli, Aurora, et al.
Published: (2024)
by: Micheli, Aurora, et al.
Published: (2024)
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
by: Elmaaroufi, Karim, et al.
Published: (2025)
by: Elmaaroufi, Karim, et al.
Published: (2025)
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
by: Du, Hao, et al.
Published: (2025)
by: Du, Hao, et al.
Published: (2025)
Do Object Detection Localization Errors Affect Human Performance and Trust?
by: de Witte, Sven, et al.
Published: (2024)
by: de Witte, Sven, et al.
Published: (2024)
Aligning Object Detector Bounding Boxes with Human Preference
by: Strafforello, Ombretta, et al.
Published: (2024)
by: Strafforello, Ombretta, et al.
Published: (2024)
GazeHTA: End-to-end Gaze Target Detection with Head-Target Association
by: Lin, Zhi-Yi, et al.
Published: (2024)
by: Lin, Zhi-Yi, et al.
Published: (2024)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
MSD: A Benchmark Dataset for Floor Plan Generation of Building Complexes
by: van Engelenburg, Casper, et al.
Published: (2024)
by: van Engelenburg, Casper, et al.
Published: (2024)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Agentic 3D Scene Generation with Spatially Contextualized VLMs
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Linear Scaling Video VLMs for Long Video Understanding
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
by: Eyzaguirre, Cristobal, et al.
Published: (2026)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025)
by: Elmansoury, Sary, et al.
Published: (2025)
Predictive Temporal Attention on Event-based Video Stream for Energy-efficient Situation Awareness
by: Bu, Yiming, et al.
Published: (2024)
by: Bu, Yiming, et al.
Published: (2024)
Clapper: Compact Learning and Video Representation in VLMs
by: Kong, Lingyu, et al.
Published: (2025)
by: Kong, Lingyu, et al.
Published: (2025)
Video Text Preservation with Synthetic Text-Rich Videos
by: Liu, Ziyang, et al.
Published: (2025)
by: Liu, Ziyang, et al.
Published: (2025)
Pushing Joint Image Denoising and Classification to the Edge
by: Markhorst, Thomas C, et al.
Published: (2024)
by: Markhorst, Thomas C, et al.
Published: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
Preserving Localized Patch Semantics in VLMs
by: Esmaeilkhani, Parsa, et al.
Published: (2026)
by: Esmaeilkhani, Parsa, et al.
Published: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Similar Items
-
Evaluation of Vision-LLMs in Surveillance Video
by: Benschop, Pascal, et al.
Published: (2025) -
Identifying Ethical Biases in Action Recognition Models
by: Baltaretu, Ana, et al.
Published: (2026) -
HAVANA: Hierarchical stochastic neighbor embedding for Accelerated Video ANnotAtions
by: Bobe, Alexandru, et al.
Published: (2024) -
Compositional Scene Understanding through Inverse Generative Modeling
by: Wang, Yanbo, et al.
Published: (2025) -
Slot-VAE: Object-Centric Scene Generation with Slot Attention
by: Wang, Yanbo, et al.
Published: (2023)