SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jintao, Huang, Haofeng, Zhang, Pengle, Wei, Jia, Zhu, Jun, Chen, Jianfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
PORTAL: Controllable Landscape Generator for Continuous Optimization-Part I: Framework
by: Yazdani, Danial, et al.
Published: (2025)
by: Yazdani, Danial, et al.
Published: (2025)
Using Evolutionary Algorithms to Find Cache-Friendly Generalized Morton Layouts for Arrays
by: Swatman, Stephen Nicholas, et al.
Published: (2023)
by: Swatman, Stephen Nicholas, et al.
Published: (2023)
Time-Fair Benchmarking for Metaheuristics: A Restart-Fair Protocol for Fixed-Time Comparisons
by: Lian, Junbo Jacob
Published: (2025)
by: Lian, Junbo Jacob
Published: (2025)
Explicit Sign-Magnitude Encoders Enable Power-Efficient Multipliers
by: Arnold, Felix, et al.
Published: (2025)
by: Arnold, Felix, et al.
Published: (2025)
Performance Evaluation of Bitstring Representations in a Linear Genetic Programming Framework
by: Meli, Clyde, et al.
Published: (2025)
by: Meli, Clyde, et al.
Published: (2025)
Fast Algorithms for Spiking Neural Network Simulation with FPGAs
by: Lindqvist, Björn A., et al.
Published: (2024)
by: Lindqvist, Björn A., et al.
Published: (2024)
HRA: A Multi-Criteria Framework for Ranking Metaheuristic Optimization Algorithms
by: Goula, Evgenia-Maria K., et al.
Published: (2024)
by: Goula, Evgenia-Maria K., et al.
Published: (2024)
A Performance Analysis of Basin Hopping Compared to Established Metaheuristics for Global Optimization
by: Baioletti, Marco, et al.
Published: (2024)
by: Baioletti, Marco, et al.
Published: (2024)
Multi-threaded Memory Efficient Crossover in C++ for Generational Genetic Programming
by: Langdon, W. B.
Published: (2020)
by: Langdon, W. B.
Published: (2020)
Combining Aggregated Attention and Transformer Architecture for Accurate and Efficient Performance of Spiking Neural Networks
by: Zhang, Hangming, et al.
Published: (2024)
by: Zhang, Hangming, et al.
Published: (2024)
Gated Attention Coding for Training High-performance and Efficient Spiking Neural Networks
by: Qiu, Xuerui, et al.
Published: (2023)
by: Qiu, Xuerui, et al.
Published: (2023)
SpikeAtConv: An Integrated Spiking-Convolutional Attention Architecture for Energy-Efficient Neuromorphic Vision Processing
by: Liao, Wangdan, et al.
Published: (2024)
by: Liao, Wangdan, et al.
Published: (2024)
Spiking Transformer with Spatial-Temporal Attention
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
Fast Clifford Neural Layers
by: Xia, Tianxiang, et al.
Published: (2025)
by: Xia, Tianxiang, et al.
Published: (2025)
ASPO: Constraint-Aware Bayesian Optimization for FPGA-based Soft Processors
by: Wu, Haoran, et al.
Published: (2025)
by: Wu, Haoran, et al.
Published: (2025)
Enhanced Innovized Repair Operator for Evolutionary Multi- and Many-objective Optimization
by: Mittal, Sukrit, et al.
Published: (2020)
by: Mittal, Sukrit, et al.
Published: (2020)
Efficient Brain Imaging Analysis for Alzheimer's and Dementia Detection Using Convolution-Derivative Operations
by: Mustafa, Yasmine, et al.
Published: (2024)
by: Mustafa, Yasmine, et al.
Published: (2024)
ROIDS: Robust Outlier-Aware Informed Down-Sampling
by: Geiger, Alina, et al.
Published: (2026)
by: Geiger, Alina, et al.
Published: (2026)
SpikePool: Event-driven Spiking Transformer with Pooling Attention
by: Lee, Donghyun, et al.
Published: (2025)
by: Lee, Donghyun, et al.
Published: (2025)
Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression
by: Makenali, Sara, et al.
Published: (2025)
by: Makenali, Sara, et al.
Published: (2025)
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
by: Yang, Liu, et al.
Published: (2026)
by: Yang, Liu, et al.
Published: (2026)
Realtime Facial Expression Recognition: Neuromorphic Hardware vs. Edge AI Accelerators
by: Smith, Heath, et al.
Published: (2024)
by: Smith, Heath, et al.
Published: (2024)
Attention to task structure for cognitive flexibility
by: Zhang, Xiaoyu K., et al.
Published: (2026)
by: Zhang, Xiaoyu K., et al.
Published: (2026)
Neural Dynamics Self-Attention for Spiking Transformers
by: Zhang, Dehao, et al.
Published: (2026)
by: Zhang, Dehao, et al.
Published: (2026)
Breaking Global Self-Attention Bottlenecks in Transformer-based Spiking Neural Networks with Local Structure-Aware Self-Attention
by: Li, Lingdong, et al.
Published: (2026)
by: Li, Lingdong, et al.
Published: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
An Efficient Evolutionary Algorithm for Few-for-Many Optimization
by: Shang, Ke, et al.
Published: (2026)
by: Shang, Ke, et al.
Published: (2026)
Attention-Driven LPLC2 Neural Ensemble Model for Multi-Target Looming Detection and Localization
by: Liu, Renyuan, et al.
Published: (2025)
by: Liu, Renyuan, et al.
Published: (2025)
Enhancing Adaptive History Reserving by Spiking Convolutional Block Attention Module in Recurrent Neural Networks
by: Xu, Qi, et al.
Published: (2024)
by: Xu, Qi, et al.
Published: (2024)
Attention-based UNet enabled Lightweight Image Semantic Communication System over Internet of Things
by: Ma, Guoxin, et al.
Published: (2024)
by: Ma, Guoxin, et al.
Published: (2024)
Efficient Multiplayer Battle Game Optimizer for Adversarial Robust Neural Architecture Search
by: Zhong, Rui, et al.
Published: (2024)
by: Zhong, Rui, et al.
Published: (2024)
Integer-State Dynamics of Quantized Spiking Neural Networks for Efficient Hardware Acceleration
by: Zhang, Lei
Published: (2026)
by: Zhang, Lei
Published: (2026)
Quantization Meets Spikes: Lossless Conversion in the First Timestep via Polarity Multi-Spike Mapping
by: Zhang, Hangming, et al.
Published: (2025)
by: Zhang, Hangming, et al.
Published: (2025)
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
by: Huang, Yulong, et al.
Published: (2026)
by: Huang, Yulong, et al.
Published: (2026)
Twin Network Augmentation: A Novel Training Strategy for Improved Spiking Neural Networks and Efficient Weight Quantization
by: Deckers, Lucas, et al.
Published: (2024)
by: Deckers, Lucas, et al.
Published: (2024)
Effective Self-Attention-Based Deep Learning Model with Evolutionary Grid Search for Robust Wave Farm Energy Forecasting
by: Dehkordi, Amin Abdollahi, et al.
Published: (2025)
by: Dehkordi, Amin Abdollahi, et al.
Published: (2025)
Exploring Extreme Quantization in Spiking Language Models
by: Bal, Malyaban, et al.
Published: (2024)
by: Bal, Malyaban, et al.
Published: (2024)
Similar Items
-
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
by: Zhang, Jintao, et al.
Published: (2025) -
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025) -
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024) -
PORTAL: Controllable Landscape Generator for Continuous Optimization-Part I: Framework
by: Yazdani, Danial, et al.
Published: (2025) -
Using Evolutionary Algorithms to Find Cache-Friendly Generalized Morton Layouts for Arrays
by: Swatman, Stephen Nicholas, et al.
Published: (2023)