LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Zhe, Lian, Sen, Wang, Changwei, Zhang, Muyang, Tan, Tianlong, Xu, Rongtao, Meng, Weiliang, Zhang, Xiaopeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913053996482560
author Feng, Zhe
Lian, Sen
Wang, Changwei
Zhang, Muyang
Tan, Tianlong
Xu, Rongtao
Meng, Weiliang
Zhang, Xiaopeng
author_facet Feng, Zhe
Lian, Sen
Wang, Changwei
Zhang, Muyang
Tan, Tianlong
Xu, Rongtao
Meng, Weiliang
Zhang, Xiaopeng
contents The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear attention variants often replace the softmax with Gaussian kernels to reduce complexity, but such approximations lack theoretical grounding and tend to oversuppress mid-range token interactions. We propose LaplacianFormer, a Transformer variant that employs a Laplacian kernel as a principled alternative to softmax, motivated by empirical observations and theoretical analysis. To address expressiveness degradation under low-rank approximations, we introduce a provably injective feature map that retains fine-grained token information. For efficient computation, we adopt a Nyström approximation of the kernel matrix and solve the resulting system using Newton--Schulz iteration, avoiding costly matrix inversion and SVD. We further develop custom CUDA implementations for both the kernel and solver, enabling high-throughput forward and backward passes suitable for edge deployment. Experiments on ImageNet show that LaplacianFormer achieves strong performance-efficiency trade-offs while improving attention expressiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20368
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
Feng, Zhe
Lian, Sen
Wang, Changwei
Zhang, Muyang
Tan, Tianlong
Xu, Rongtao
Meng, Weiliang
Zhang, Xiaopeng
Computer Vision and Pattern Recognition
Artificial Intelligence
The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear attention variants often replace the softmax with Gaussian kernels to reduce complexity, but such approximations lack theoretical grounding and tend to oversuppress mid-range token interactions. We propose LaplacianFormer, a Transformer variant that employs a Laplacian kernel as a principled alternative to softmax, motivated by empirical observations and theoretical analysis. To address expressiveness degradation under low-rank approximations, we introduce a provably injective feature map that retains fine-grained token information. For efficient computation, we adopt a Nyström approximation of the kernel matrix and solve the resulting system using Newton--Schulz iteration, avoiding costly matrix inversion and SVD. We further develop custom CUDA implementations for both the kernel and solver, enabling high-throughput forward and backward passes suitable for edge deployment. Experiments on ImageNet show that LaplacianFormer achieves strong performance-efficiency trade-offs while improving attention expressiveness.
title LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.20368