Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911599469527040 |
|---|---|
| author | Zuhri, Zayd M. K. Fuadi, Erland Hilman Aji, Alham Fikri |
| author_facet | Zuhri, Zayd M. K. Fuadi, Erland Hilman Aji, Alham Fikri |
| contents | We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M and 1.8B parameter models demonstrate that softpick achieves 0\% sink rate consistently. The softpick transformers produce hidden states with significantly lower kurtosis and creates sparse attention maps. Quantized models using softpick outperform softmax on standard benchmarks, with a particularly pronounced advantage at lower bit precisions. Our analysis and discussion shows how softpick has the potential to open new possibilities for quantization, low-precision training, sparsity optimization, pruning, and interpretability. Our code: https://github.com/zaydzuhri/softpick-attention |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_20966 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Softpick: No Attention Sink, No Massive Activations with Rectified Softmax Zuhri, Zayd M. K. Fuadi, Erland Hilman Aji, Alham Fikri Machine Learning We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M and 1.8B parameter models demonstrate that softpick achieves 0\% sink rate consistently. The softpick transformers produce hidden states with significantly lower kurtosis and creates sparse attention maps. Quantized models using softpick outperform softmax on standard benchmarks, with a particularly pronounced advantage at lower bit precisions. Our analysis and discussion shows how softpick has the potential to open new possibilities for quantization, low-precision training, sparsity optimization, pruning, and interpretability. Our code: https://github.com/zaydzuhri/softpick-attention |
| title | Softpick: No Attention Sink, No Massive Activations with Rectified Softmax |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2504.20966 |