SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luna, Jose Miguel, Bouhsine, Taha, Choromanski, Krzysztof
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908822257270784
author Luna, Jose Miguel
Bouhsine, Taha
Choromanski, Krzysztof
author_facet Luna, Jose Miguel
Bouhsine, Taha
Choromanski, Krzysztof
contents We propose a new class of linear-time attention mechanisms based on a relaxed and computationally efficient formulation of the recently introduced E-Product, often referred to as the Yat-kernel (Bouhsine, 2025). The resulting interactions are geometry-aware and inspired by inverse-square interactions in physics. Our method, Spherical Linearized Attention with Yat Kernels (SLAY), constrains queries and keys to the unit sphere so that attention depends only on angular alignment. Using Bernstein's theorem, we express the spherical Yat-kernel as a nonnegative mixture of polynomial-exponential product kernels and derive a strictly positive random-feature approximation enabling linear-time O(L) attention. We establish positive definiteness and boundedness on the sphere and show that the estimator yields well-defined, nonnegative attention scores. Empirically, SLAY achieves performance that is nearly indistinguishable from standard softmax attention while retaining linear time and memory scaling, and consistently outperforms prior linear-time attention mechanisms such as Performers and Cosformers. To the best of our knowledge, SLAY represents the closest linear-time approximation to softmax attention reported to date, enabling scalable Transformers without the typical performance trade-offs of attention linearization.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04915
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel
Luna, Jose Miguel
Bouhsine, Taha
Choromanski, Krzysztof
Machine Learning
Artificial Intelligence
I.2.6 Learning
We propose a new class of linear-time attention mechanisms based on a relaxed and computationally efficient formulation of the recently introduced E-Product, often referred to as the Yat-kernel (Bouhsine, 2025). The resulting interactions are geometry-aware and inspired by inverse-square interactions in physics. Our method, Spherical Linearized Attention with Yat Kernels (SLAY), constrains queries and keys to the unit sphere so that attention depends only on angular alignment. Using Bernstein's theorem, we express the spherical Yat-kernel as a nonnegative mixture of polynomial-exponential product kernels and derive a strictly positive random-feature approximation enabling linear-time O(L) attention. We establish positive definiteness and boundedness on the sphere and show that the estimator yields well-defined, nonnegative attention scores. Empirically, SLAY achieves performance that is nearly indistinguishable from standard softmax attention while retaining linear time and memory scaling, and consistently outperforms prior linear-time attention mechanisms such as Performers and Cosformers. To the best of our knowledge, SLAY represents the closest linear-time approximation to softmax attention reported to date, enabling scalable Transformers without the typical performance trade-offs of attention linearization.
title SLAY: Geometry-Aware Spherical Linearized Attention with Yat-Kernel
topic Machine Learning
Artificial Intelligence
I.2.6 Learning
url https://arxiv.org/abs/2602.04915