AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915990435004416 |
|---|---|
| author | Li, Handong Liu, Zikang Guo, Longteng Yue, Tongtian Tang, Yepeng Zhu, Xinxin Zheng, Chuanyang Wang, Ziming Wang, Zhibin Song, Jun Yu, Cheng Zheng, Bo Liu, Jing |
| author_facet | Li, Handong Liu, Zikang Guo, Longteng Yue, Tongtian Tang, Yepeng Zhu, Xinxin Zheng, Chuanyang Wang, Ziming Wang, Zhibin Song, Jun Yu, Cheng Zheng, Bo Liu, Jing |
| contents | Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This paper introduces AdaSpark, an adaptive sparsity framework designed to address these limitations. AdaSpark first partitions video inputs into 3D spatio-temporal cubes. It then employs two co-designed, context-aware components: (1) Adaptive Cube-Selective Attention (AdaS-Attn), which adaptively selects a subset of relevant video cubes to attend for each query token, and (2) Adaptive Token-Selective FFN (AdaS-FFN), which selectively processes only the most salient tokens within each cube. An entropy-based (Top-p) selection mechanism adaptively allocates computational resources based on input complexity. Experiments demonstrate that AdaSpark significantly reduces computational load by up to 57% FLOPs while maintaining comparable performance to dense models and preserving fine-grained, long-range dependencies, as validated on challenging hour-scale video benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_08077 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding Li, Handong Liu, Zikang Guo, Longteng Yue, Tongtian Tang, Yepeng Zhu, Xinxin Zheng, Chuanyang Wang, Ziming Wang, Zhibin Song, Jun Yu, Cheng Zheng, Bo Liu, Jing Computer Vision and Pattern Recognition Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fine-grained perception through irreversible information disposal or inhibit long-range temporal modeling via rigid, predefined sparse patterns. This paper introduces AdaSpark, an adaptive sparsity framework designed to address these limitations. AdaSpark first partitions video inputs into 3D spatio-temporal cubes. It then employs two co-designed, context-aware components: (1) Adaptive Cube-Selective Attention (AdaS-Attn), which adaptively selects a subset of relevant video cubes to attend for each query token, and (2) Adaptive Token-Selective FFN (AdaS-FFN), which selectively processes only the most salient tokens within each cube. An entropy-based (Top-p) selection mechanism adaptively allocates computational resources based on input complexity. Experiments demonstrate that AdaSpark significantly reduces computational load by up to 57% FLOPs while maintaining comparable performance to dense models and preserving fine-grained, long-range dependencies, as validated on challenging hour-scale video benchmarks. |
| title | AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.08077 |