Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuhri, Zayd M. K., Fuadi, Erland Hilman, Aji, Alham Fikri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911599469527040
author Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
author_facet Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
contents We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M and 1.8B parameter models demonstrate that softpick achieves 0\% sink rate consistently. The softpick transformers produce hidden states with significantly lower kurtosis and creates sparse attention maps. Quantized models using softpick outperform softmax on standard benchmarks, with a particularly pronounced advantage at lower bit precisions. Our analysis and discussion shows how softpick has the potential to open new possibilities for quantization, low-precision training, sparsity optimization, pruning, and interpretability. Our code: https://github.com/zaydzuhri/softpick-attention
format Preprint
id arxiv_https___arxiv_org_abs_2504_20966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
Zuhri, Zayd M. K.
Fuadi, Erland Hilman
Aji, Alham Fikri
Machine Learning
We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention sink and massive activations. Our experiments with 340M and 1.8B parameter models demonstrate that softpick achieves 0\% sink rate consistently. The softpick transformers produce hidden states with significantly lower kurtosis and creates sparse attention maps. Quantized models using softpick outperform softmax on standard benchmarks, with a particularly pronounced advantage at lower bit precisions. Our analysis and discussion shows how softpick has the potential to open new possibilities for quantization, low-precision training, sparsity optimization, pruning, and interpretability. Our code: https://github.com/zaydzuhri/softpick-attention
title Softpick: No Attention Sink, No Massive Activations with Rectified Softmax
topic Machine Learning
url https://arxiv.org/abs/2504.20966