Sliced ReLU attention: Quasi-linear contextual expressivity via sorting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vialard, François-Xavier, Boufadène, Siwan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910011451506688
author Vialard, François-Xavier
Boufadène, Siwan
author_facet Vialard, François-Xavier
Boufadène, Siwan
contents We introduce sliced ReLU attention, a new attention mechanism that departs structurally from both softmax and its approximation alternatives. Instead of applying a nonlinearity to pairwise dot products, we operate on one-dimensional projections of key--query differences and leverage sorting to obtain quasi-linear complexity. This construction yields a differentiable, non-symmetric kernel that can be computed in O(n log(n)) through a sorting procedure, making it suitable for very long contexts. Beyond computational benefits, the model retains strong theoretical expressive power: we establish two in-context expressivity results, previously known for softmax attention, showing that sliced ReLU attention preserves the ability to perform nontrivial sequence-to-sequence disentangling tasks and satisfies a contextual universal approximation property. Finally, we illustrate the potential practical interest of this kernel in small to medium-scale experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
Vialard, François-Xavier
Boufadène, Siwan
Machine Learning
We introduce sliced ReLU attention, a new attention mechanism that departs structurally from both softmax and its approximation alternatives. Instead of applying a nonlinearity to pairwise dot products, we operate on one-dimensional projections of key--query differences and leverage sorting to obtain quasi-linear complexity. This construction yields a differentiable, non-symmetric kernel that can be computed in O(n log(n)) through a sorting procedure, making it suitable for very long contexts. Beyond computational benefits, the model retains strong theoretical expressive power: we establish two in-context expressivity results, previously known for softmax attention, showing that sliced ReLU attention preserves the ability to perform nontrivial sequence-to-sequence disentangling tasks and satisfies a contextual universal approximation property. Finally, we illustrate the potential practical interest of this kernel in small to medium-scale experiments.
title Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
topic Machine Learning
url https://arxiv.org/abs/2512.11411