A Survey of Video Datasets for Grounded Event Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Sanders, Kate, Van Durme, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
by: Sanders, Kate, et al.
Published: (2025)
by: Sanders, Kate, et al.
Published: (2025)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
Grounding Partially-Defined Events in Multimodal Data
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
by: Guo, Zhenghui, et al.
Published: (2026)
by: Guo, Zhenghui, et al.
Published: (2026)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
by: Chen, Ruizhe, et al.
Published: (2025)
by: Chen, Ruizhe, et al.
Published: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
by: Singh, Shivam, et al.
Published: (2026)
by: Singh, Shivam, et al.
Published: (2026)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
by: Cheng, Dingxin, et al.
Published: (2024)
by: Cheng, Dingxin, et al.
Published: (2024)
Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
by: Mago, Gowreesh, et al.
Published: (2025)
by: Mago, Gowreesh, et al.
Published: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
by: Akter, Sanjeda, et al.
Published: (2025)
by: Akter, Sanjeda, et al.
Published: (2025)
PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
by: Zhou, Zhiyu, et al.
Published: (2026)
by: Zhou, Zhiyu, et al.
Published: (2026)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
by: Zhou, Pengyuan, et al.
Published: (2024)
by: Zhou, Pengyuan, et al.
Published: (2024)
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
by: Kim, Sunoh, et al.
Published: (2023)
by: Kim, Sunoh, et al.
Published: (2023)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
by: Li, Shicheng, et al.
Published: (2023)
by: Li, Shicheng, et al.
Published: (2023)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
by: Vu, Huu-An, et al.
Published: (2025)
by: Vu, Huu-An, et al.
Published: (2025)
Transformer-based Spatial Grounding: A Comprehensive Survey
by: Haq, Ijazul, et al.
Published: (2025)
by: Haq, Ijazul, et al.
Published: (2025)
A Survey of 3D Reconstruction with Event Cameras
by: Xu, Chuanzhi, et al.
Published: (2025)
by: Xu, Chuanzhi, et al.
Published: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
Video Panels for Long Video Understanding
by: Doorenbos, Lars, et al.
Published: (2025)
by: Doorenbos, Lars, et al.
Published: (2025)
WikiVideo: Article Generation from Multiple Videos
by: Martin, Alexander, et al.
Published: (2025)
by: Martin, Alexander, et al.
Published: (2025)
EventVL: Understand Event Streams via Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
When and Where do Events Switch in Multi-Event Video Generation?
by: Liao, Ruotong, et al.
Published: (2025)
by: Liao, Ruotong, et al.
Published: (2025)
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024)
by: Zhang, Gengyuan, et al.
Published: (2024)
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Similar Items
-
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
by: Sanders, Kate, et al.
Published: (2025) -
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
by: Sanders, Kate, et al.
Published: (2024) -
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025) -
Grounding Partially-Defined Events in Multimodal Data
by: Sanders, Kate, et al.
Published: (2024) -
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)