EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Wenhao, Dong, Xin, Li, Yue, Shi, Haoyuan, Xiong, Zhiwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914168736579584
author Xu, Wenhao
Dong, Xin
Li, Yue
Shi, Haoyuan
Xiong, Zhiwei
author_facet Xu, Wenhao
Dong, Xin
Li, Yue
Shi, Haoyuan
Xiong, Zhiwei
contents Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspired by event-based vision, we propose an event-guided, training-free framework for efficient spatio-temporal understanding, named EventSTU. In the temporal domain, we design a coarse-to-fine keyframe sampling algorithm that exploits the change-triggered property of event cameras to eliminate redundant frames. In the spatial domain, we design an adaptive token pruning algorithm that leverages the visual saliency of events as a zero-cost prior to guide spatial reduction. From a holistic spatio-temporal perspective, we further integrate question relevance from keyframe sampling to adaptively allocate token pruning budgets. To facilitate evaluation, we construct EventBench, the first event-inclusive, human-annotated multimodal benchmark that covers diverse real-world scenarios. Beyond physical event cameras, EventSTU also supports general video understanding using simulated events. Comprehensive experiments show that EventSTU achieves 3.01x FLOPs reduction and 3.10x prefilling speedup over the strongest baseline while still improving performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
Xu, Wenhao
Dong, Xin
Li, Yue
Shi, Haoyuan
Xiong, Zhiwei
Computer Vision and Pattern Recognition
Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspired by event-based vision, we propose an event-guided, training-free framework for efficient spatio-temporal understanding, named EventSTU. In the temporal domain, we design a coarse-to-fine keyframe sampling algorithm that exploits the change-triggered property of event cameras to eliminate redundant frames. In the spatial domain, we design an adaptive token pruning algorithm that leverages the visual saliency of events as a zero-cost prior to guide spatial reduction. From a holistic spatio-temporal perspective, we further integrate question relevance from keyframe sampling to adaptively allocate token pruning budgets. To facilitate evaluation, we construct EventBench, the first event-inclusive, human-annotated multimodal benchmark that covers diverse real-world scenarios. Beyond physical event cameras, EventSTU also supports general video understanding using simulated events. Comprehensive experiments show that EventSTU achieves 3.01x FLOPs reduction and 3.10x prefilling speedup over the strongest baseline while still improving performance.
title EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18920