EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Wenhao, Dong, Xin, Li, Yue, Shi, Haoyuan, Xiong, Zhiwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
by: Zhou, Hanyu, et al.
Published: (2024)
by: Zhou, Hanyu, et al.
Published: (2024)
EventGPT: Event Stream Understanding with Multimodal Large Language Models
by: Liu, Shaoyu, et al.
Published: (2024)
by: Liu, Shaoyu, et al.
Published: (2024)
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
by: Ge, Chengjie, et al.
Published: (2025)
by: Ge, Chengjie, et al.
Published: (2025)
EventVL: Understand Event Streams via Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
Event-boosted Deformable 3D Gaussians for Dynamic Scene Reconstruction
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Spatio-Temporal State Space Model For Efficient Event-Based Optical Flow
by: Humais, Muhammad Ahmed, et al.
Published: (2025)
by: Humais, Muhammad Ahmed, et al.
Published: (2025)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
by: Gao, Shida, et al.
Published: (2025)
by: Gao, Shida, et al.
Published: (2025)
Event-assisted Low-Light Video Object Segmentation
by: Li, Hebei, et al.
Published: (2024)
by: Li, Hebei, et al.
Published: (2024)
Context-Guided Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
Identifying Spatio-Temporal Drivers of Extreme Events
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
by: Eddin, Mohamad Hakam Shams, et al.
Published: (2024)
Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
by: Li, Renkai, et al.
Published: (2025)
by: Li, Renkai, et al.
Published: (2025)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
by: Zhuang, Weijun, et al.
Published: (2026)
by: Zhuang, Weijun, et al.
Published: (2026)
Event-to-Video Reconstruction using Spatio-Temporal and Frequency-Enhanced Deep Neural Networks
by: Maqsood, Ramna, et al.
Published: (2026)
by: Maqsood, Ramna, et al.
Published: (2026)
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
by: Lin, Xinying, et al.
Published: (2026)
by: Lin, Xinying, et al.
Published: (2026)
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
EGVD: Event-Guided Video Diffusion Model for Physically Realistic Large-Motion Frame Interpolation
by: Zhang, Ziran, et al.
Published: (2025)
by: Zhang, Ziran, et al.
Published: (2025)
Nonlinear Motion-Guided and Spatio-Temporal Aware Network for Unsupervised Event-Based Optical Flow
by: Liu, Zuntao, et al.
Published: (2025)
by: Liu, Zuntao, et al.
Published: (2025)
Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
by: Zhu, Lin, et al.
Published: (2024)
by: Zhu, Lin, et al.
Published: (2024)
CompEvent: Complex-valued Event-RGB Fusion for Low-light Video Enhancement and Deblurring
by: Zhong, Mingchen, et al.
Published: (2025)
by: Zhong, Mingchen, et al.
Published: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
by: Zeng, Xiangyu, et al.
Published: (2025)
by: Zeng, Xiangyu, et al.
Published: (2025)
From Events to Clarity: The Event-Guided Diffusion Framework for Dehazing
by: Wang, Ling, et al.
Published: (2025)
by: Wang, Ling, et al.
Published: (2025)
Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
by: Lu, Jinpeng, et al.
Published: (2025)
by: Lu, Jinpeng, et al.
Published: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
by: Shi, Jiapeng, et al.
Published: (2026)
by: Shi, Jiapeng, et al.
Published: (2026)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
Event-Enhanced Blurry Video Super-Resolution
by: Kai, Dachun, et al.
Published: (2025)
by: Kai, Dachun, et al.
Published: (2025)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
by: Qin, Haotong, et al.
Published: (2025)
by: Qin, Haotong, et al.
Published: (2025)
EventAug: Multifaceted Spatio-Temporal Data Augmentation Methods for Event-based Learning
by: Tian, Yukun, et al.
Published: (2024)
by: Tian, Yukun, et al.
Published: (2024)
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
by: Yan, Pengyu, et al.
Published: (2026)
by: Yan, Pengyu, et al.
Published: (2026)
Temporal-Guided Visual Foundation Models for Event-Based Vision
by: Xia, Ruihao, et al.
Published: (2025)
by: Xia, Ruihao, et al.
Published: (2025)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data
by: Gu, Cheng, et al.
Published: (2021)
by: Gu, Cheng, et al.
Published: (2021)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
by: An, Hongyu, et al.
Published: (2024)
by: An, Hongyu, et al.
Published: (2024)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
by: Ma, Jingtian, et al.
Published: (2025)
by: Ma, Jingtian, et al.
Published: (2025)
ScVLM: Enhancing Vision-Language Model for Safety-Critical Event Understanding
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
Similar Items
-
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
by: Xu, Wenhao, et al.
Published: (2024) -
JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
by: Zhou, Hanyu, et al.
Published: (2024) -
EventGPT: Event Stream Understanding with Multimodal Large Language Models
by: Liu, Shaoyu, et al.
Published: (2024) -
EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
by: Ge, Chengjie, et al.
Published: (2025) -
EventVL: Understand Event Streams via Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)