SageBwd: A Trainable Low-bit Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jintao, Chen, Marco, Wang, Haoxu, Jiang, Kai, Stoica, Ion, Gonzalez, Joseph E., Chen, Jianfei, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SageAttention2++: A More Efficient Implementation of SageAttention2
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
Post-Training Sparse Attention with Double Sparsity
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
HashAttention: Semantic Sparsity for Faster Inference
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
von: Desai, Aditya, et al.
Veröffentlicht: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
von: Mishra, Mayank, et al.
Veröffentlicht: (2026)
Identifying Sensitive Weights via Post-quantization Integral
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
von: Zhu, Alan, et al.
Veröffentlicht: (2025)
Fairness in Serving Large Language Models
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
Low-bit Model Quantization for Deep Neural Networks: A Survey
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
von: Cao, Shiyi, et al.
Veröffentlicht: (2026)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
von: Lapautre, Nicolas, et al.
Veröffentlicht: (2025)
A Statistical Framework for Ranking LLM-Based Chatbots
von: Ameli, Siavash, et al.
Veröffentlicht: (2024)
von: Ameli, Siavash, et al.
Veröffentlicht: (2024)
Visual Generation Without Guidance
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
von: Chen, Huayu, et al.
Veröffentlicht: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
von: Gong, Ping, et al.
Veröffentlicht: (2025)
von: Gong, Ping, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
S*: Test Time Scaling for Code Generation
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
von: Yan, Xianglong, et al.
Veröffentlicht: (2026)
MPC-Minimized Secure LLM Inference
von: Rathee, Deevashwer, et al.
Veröffentlicht: (2024)
von: Rathee, Deevashwer, et al.
Veröffentlicht: (2024)
When Bias Meets Trainability: Connecting Theories of Initialization
von: Bassi, Alberto, et al.
Veröffentlicht: (2025)
von: Bassi, Alberto, et al.
Veröffentlicht: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
von: Jiang, Xuanlin, et al.
Veröffentlicht: (2024)
SparseDM: Toward Sparse Efficient Diffusion Models
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
von: Wang, Kafeng, et al.
Veröffentlicht: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
von: Chen, Yuxiang, et al.
Veröffentlicht: (2025)
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SLA2: Sparse-Linear Attention with Learnable Routing and QAT
von: Zhang, Jintao, et al.
Veröffentlicht: (2026) -
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2025) -
SageAttention2++: A More Efficient Implementation of SageAttention2
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)