VideoNSA: Native Sparse Attention Scales Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Enxin, Chai, Wenhao, Yang, Shusheng, Armand, Ethan, Shan, Xiaojun, Xu, Haiyang, Xie, Jianwen, Tu, Zhuowen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915761524572160
author Song, Enxin
Chai, Wenhao
Yang, Shusheng
Armand, Ethan
Shan, Xiaojun
Xu, Haiyang
Xie, Jianwen
Tu, Zhuowen
author_facet Song, Enxin
Chai, Wenhao
Yang, Shusheng
Armand, Ethan
Shan, Xiaojun
Xu, Haiyang
Xie, Jianwen
Tu, Zhuowen
contents Video understanding in multimodal language models remains limited by context length: models often miss key transition frames and struggle to maintain coherence across long time scales. To address this, we adapt Native Sparse Attention (NSA) to video-language models. Our method, VideoNSA, adapts Qwen2.5-VL through end-to-end training on a 216K video instruction dataset. We employ a hardware-aware hybrid approach to attention, preserving dense attention for text, while employing NSA for video. Compared to token-compression and training-free sparse baselines, VideoNSA achieves improved performance on long-video understanding, temporal reasoning, and spatial benchmarks. Further ablation analysis reveals four key findings: (1) reliable scaling to 128K tokens; (2) an optimal global-local attention allocation at a fixed budget; (3) task-dependent branch usage patterns; and (4) the learnable combined sparse attention help induce dynamic attention sinks. Project Page: https://enxinsong.com/VideoNSA-web/, Code: https://github.com/Espere-1119-Song/VideoNSA
format Preprint
id arxiv_https___arxiv_org_abs_2510_02295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoNSA: Native Sparse Attention Scales Video Understanding
Song, Enxin
Chai, Wenhao
Yang, Shusheng
Armand, Ethan
Shan, Xiaojun
Xu, Haiyang
Xie, Jianwen
Tu, Zhuowen
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Video understanding in multimodal language models remains limited by context length: models often miss key transition frames and struggle to maintain coherence across long time scales. To address this, we adapt Native Sparse Attention (NSA) to video-language models. Our method, VideoNSA, adapts Qwen2.5-VL through end-to-end training on a 216K video instruction dataset. We employ a hardware-aware hybrid approach to attention, preserving dense attention for text, while employing NSA for video. Compared to token-compression and training-free sparse baselines, VideoNSA achieves improved performance on long-video understanding, temporal reasoning, and spatial benchmarks. Further ablation analysis reveals four key findings: (1) reliable scaling to 128K tokens; (2) an optimal global-local attention allocation at a fixed budget; (3) task-dependent branch usage patterns; and (4) the learnable combined sparse attention help induce dynamic attention sinks. Project Page: https://enxinsong.com/VideoNSA-web/, Code: https://github.com/Espere-1119-Song/VideoNSA
title VideoNSA: Native Sparse Attention Scales Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.02295