Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yanke, Li, Yiduo, Tang, Hanlin, Li, Maohua, Liu, Kan, Tao, Lan, Qu, Lin, Yao, Yuan, Ma, Xiaoxing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Rethinking Cross-Layer Information Routing in Diffusion Transformers
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
by: Lan, Libin, et al.
Published: (2025)
by: Lan, Libin, et al.
Published: (2025)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
Native 3D Editing with Full Attention
by: Cai, Weiwei, et al.
Published: (2025)
by: Cai, Weiwei, et al.
Published: (2025)
RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models
by: Cong, Xing, et al.
Published: (2026)
by: Cong, Xing, et al.
Published: (2026)
Attention-Based Reconstruction of Full-Field Tsunami Waves from Sparse Tsunameter Networks
by: McDugald, Edward, et al.
Published: (2024)
by: McDugald, Edward, et al.
Published: (2024)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
by: Huang, Yufeng
Published: (2026)
by: Huang, Yufeng
Published: (2026)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Task Abstention for Large Language Models in Code Generation
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
Full-Body Motion Reconstruction with Sparse Sensing from Graph Perspective
by: Yao, Feiyu, et al.
Published: (2024)
by: Yao, Feiyu, et al.
Published: (2024)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
by: Xu, Zhihong, et al.
Published: (2025)
by: Xu, Zhihong, et al.
Published: (2025)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
Rethinking Transformer Connectivity: TLinFormer, A Path to Exact, Full Context-Aware Linear Attention
by: Tang, Zhongpan
Published: (2025)
by: Tang, Zhongpan
Published: (2025)
YOLO-PRO: Enhancing Instance-Specific Object Detection with Full-Channel Global Self-Attention
by: Huang, Lin, et al.
Published: (2025)
by: Huang, Lin, et al.
Published: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
by: Ye, Xiaowei, et al.
Published: (2026)
by: Ye, Xiaowei, et al.
Published: (2026)
Hybrid Focal and Full-Range Attention Based Graph Transformers
by: Zhu, Minhong, et al.
Published: (2023)
by: Zhu, Minhong, et al.
Published: (2023)
Bidirectional Sparse Attention for Faster Video Diffusion Training
by: Zhan, Chenlu, et al.
Published: (2025)
by: Zhan, Chenlu, et al.
Published: (2025)
Thin Keys, Full Values: Reducing KV Cache via Low-Dimensional Attention Selection
by: Yao, Hengshuai, et al.
Published: (2026)
by: Yao, Hengshuai, et al.
Published: (2026)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
by: Lin, Jiawei, et al.
Published: (2025)
by: Lin, Jiawei, et al.
Published: (2025)
Adapters Strike Back
by: Steitz, Jan-Martin O., et al.
Published: (2024)
by: Steitz, Jan-Martin O., et al.
Published: (2024)
Minimax Strikes Back
by: Cohen-Solal, Quentin, et al.
Published: (2020)
by: Cohen-Solal, Quentin, et al.
Published: (2020)
The Oval Strikes Back
by: Di Giusto, Andrea, et al.
Published: (2026)
by: Di Giusto, Andrea, et al.
Published: (2026)
BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation
by: Gu, Youping, et al.
Published: (2025)
by: Gu, Youping, et al.
Published: (2025)
Fair Conformal Classification via Learning Representation-Based Groups
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
LVSA: Training-Free Sparse Attention for Long Video Diffusion
by: Glorian, Gael, et al.
Published: (2026)
by: Glorian, Gael, et al.
Published: (2026)
Sparse Full Configuration Interaction
by: Wang, Lijun
Published: (2023)
by: Wang, Lijun
Published: (2023)
Kryptonite-N: Machine Learning Strikes Back
by: Li, Albus, et al.
Published: (2024)
by: Li, Albus, et al.
Published: (2024)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
by: Wang, Huihan, et al.
Published: (2025)
by: Wang, Huihan, et al.
Published: (2025)
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
by: Wang, Hanxiao, et al.
Published: (2025)
by: Wang, Hanxiao, et al.
Published: (2025)
Depth-Structured Music Recurrence: Budgeted Recurrent Attention for Full-Piece Symbolic Music Modeling
by: Yi, Yungang, et al.
Published: (2026)
by: Yi, Yungang, et al.
Published: (2026)
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning
by: Xu, Huimin, et al.
Published: (2025)
by: Xu, Huimin, et al.
Published: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
TCSAFormer : Efficient Vision Transformer With Token Compression and Sparse Attention for Medical Image Segmentation
by: Zunhui Xia, et al.
Published: (2026)
by: Zunhui Xia, et al.
Published: (2026)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Similar Items
-
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025) -
Rethinking Cross-Layer Information Routing in Diffusion Transformers
by: Xu, Chao, et al.
Published: (2026) -
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
by: Lan, Libin, et al.
Published: (2025) -
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025) -
Native 3D Editing with Full Attention
by: Cai, Weiwei, et al.
Published: (2025)