PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiaolong, Gu, Youping, Lin, Xi, Wang, Weijie, Zhuang, Bohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
von: Gu, Youping, et al.
Veröffentlicht: (2025)
von: Gu, Youping, et al.
Veröffentlicht: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
von: Song, Enxin, et al.
Veröffentlicht: (2025)
von: Song, Enxin, et al.
Veröffentlicht: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
von: Hassani, Ali, et al.
Veröffentlicht: (2025)
von: Hassani, Ali, et al.
Veröffentlicht: (2025)
Pyramidal Flow Matching for Efficient Video Generative Modeling
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
von: Zou, Shihao, et al.
Veröffentlicht: (2025)
von: Zou, Shihao, et al.
Veröffentlicht: (2025)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction
von: Qin, Meng'en, et al.
Veröffentlicht: (2026)
von: Qin, Meng'en, et al.
Veröffentlicht: (2026)
Efficient Image Generation with Variadic Attention Heads
von: Walton, Steven, et al.
Veröffentlicht: (2022)
von: Walton, Steven, et al.
Veröffentlicht: (2022)
Adaptive Keyframe Sampling for Long Video Understanding
von: Tang, Xi, et al.
Veröffentlicht: (2025)
von: Tang, Xi, et al.
Veröffentlicht: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
von: Li, Quanhao, et al.
Veröffentlicht: (2025)
Do Language Models Understand Time?
von: Ding, Xi, et al.
Veröffentlicht: (2024)
von: Ding, Xi, et al.
Veröffentlicht: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
HourVideo: 1-Hour Video-Language Understanding
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
Motion meets Attention: Video Motion Prompts
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
BSA: Ball Sparse Attention for Large-scale Geometries
von: Brita, Catalin E., et al.
Veröffentlicht: (2025)
von: Brita, Catalin E., et al.
Veröffentlicht: (2025)
Consistent Flow Distillation for Text-to-3D Generation
von: Yan, Runjie, et al.
Veröffentlicht: (2025)
von: Yan, Runjie, et al.
Veröffentlicht: (2025)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
Point Cloud Understanding via Attention-Driven Contrastive Learning
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free Applications
von: Hu, Zixuan, et al.
Veröffentlicht: (2025)
von: Hu, Zixuan, et al.
Veröffentlicht: (2025)
DMin: Scalable Training Data Influence Estimation for Diffusion Models
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
Video Understanding by Design: How Datasets Shape Architectures and Insights
von: Wang, Lei, et al.
Veröffentlicht: (2025)
von: Wang, Lei, et al.
Veröffentlicht: (2025)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
von: Picard, David, et al.
Veröffentlicht: (2024)
von: Picard, David, et al.
Veröffentlicht: (2024)
Dual Diffusion for Unified Image Generation and Understanding
von: Li, Zijie, et al.
Veröffentlicht: (2024)
von: Li, Zijie, et al.
Veröffentlicht: (2024)
Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian Processes
von: Chen, Yingyi, et al.
Veröffentlicht: (2024)
von: Chen, Yingyi, et al.
Veröffentlicht: (2024)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
von: Susladkar, Onkar, et al.
Veröffentlicht: (2026)
von: Susladkar, Onkar, et al.
Veröffentlicht: (2026)
VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models
von: Li, Edward, et al.
Veröffentlicht: (2025)
von: Li, Edward, et al.
Veröffentlicht: (2025)
VideoPhy: Evaluating Physical Commonsense for Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
VIDEOP2R: Video Understanding from Perception to Reasoning
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
ENA: Efficient N-dimensional Attention
von: Zhong, Yibo
Veröffentlicht: (2025)
von: Zhong, Yibo
Veröffentlicht: (2025)
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
von: Du, Hongyang, et al.
Veröffentlicht: (2026)
von: Du, Hongyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
von: Gu, Youping, et al.
Veröffentlicht: (2025) -
VideoNSA: Native Sparse Attention Scales Video Understanding
von: Song, Enxin, et al.
Veröffentlicht: (2025) -
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
von: Li, Xingyang, et al.
Veröffentlicht: (2025) -
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
von: Chen, Pengtao, et al.
Veröffentlicht: (2025) -
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
von: Hassani, Ali, et al.
Veröffentlicht: (2025)