Rethinking Attention: Polynomial Alternatives to Softmax in Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saratchandran, Hemanth, Zheng, Jianqiao, Ji, Yiping, Zhang, Wenbo, Lucey, Simon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918385008246784
author Saratchandran, Hemanth
Zheng, Jianqiao
Ji, Yiping
Zhang, Wenbo
Lucey, Simon
author_facet Saratchandran, Hemanth
Zheng, Jianqiao
Ji, Yiping
Zhang, Wenbo
Lucey, Simon
contents This paper questions whether the strong performance of softmax attention in transformers stems from producing a probability distribution over inputs. Instead, we argue that softmax's effectiveness lies in its implicit regularization of the Frobenius norm of the attention matrix, which stabilizes training. Motivated by this, we explore alternative activations, specifically polynomials, that achieve a similar regularization effect. Our theoretical analysis shows that certain polynomials can serve as effective substitutes for softmax, achieving strong performance across transformer applications despite violating softmax's typical properties of positivity, normalization, and sparsity. Extensive experiments support these findings, offering a new perspective on attention mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
Saratchandran, Hemanth
Zheng, Jianqiao
Ji, Yiping
Zhang, Wenbo
Lucey, Simon
Machine Learning
Computer Vision and Pattern Recognition
This paper questions whether the strong performance of softmax attention in transformers stems from producing a probability distribution over inputs. Instead, we argue that softmax's effectiveness lies in its implicit regularization of the Frobenius norm of the attention matrix, which stabilizes training. Motivated by this, we explore alternative activations, specifically polynomials, that achieve a similar regularization effect. Our theoretical analysis shows that certain polynomials can serve as effective substitutes for softmax, achieving strong performance across transformer applications despite violating softmax's typical properties of positivity, normalization, and sparsity. Extensive experiments support these findings, offering a new perspective on attention mechanisms.
title Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.18613