Universal Approximation with Softmax Attention

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Jerry Yao-Chieh, Liu, Hude, Chen, Hong-Yu, Wu, Weimin, Liu, Han
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918249347678208
author Hu, Jerry Yao-Chieh
Liu, Hude
Chen, Hong-Yu
Wu, Weimin
Liu, Han
author_facet Hu, Jerry Yao-Chieh
Liu, Hude
Chen, Hong-Yu
Wu, Weimin
Liu, Han
contents We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on compact domains. Our main technique is a new interpolation-based method for analyzing attention's internal mechanism. This leads to our key insight: self-attention is able to approximate a generalized version of ReLU to arbitrary precision, and hence subsumes many known universal approximators. Building on these, we show that two-layer multi-head attention alone suffices as a sequence-to-sequence universal approximator. In contrast, prior works rely on feed-forward networks to establish universal approximation in Transformers. Furthermore, we extend our techniques to show that, (softmax-)attention-only layers are capable of approximating various statistical models in-context. We believe these techniques hold independent interest.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15956
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Universal Approximation with Softmax Attention
Hu, Jerry Yao-Chieh
Liu, Hude
Chen, Hong-Yu
Wu, Weimin
Liu, Han
Machine Learning
Artificial Intelligence
We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on compact domains. Our main technique is a new interpolation-based method for analyzing attention's internal mechanism. This leads to our key insight: self-attention is able to approximate a generalized version of ReLU to arbitrary precision, and hence subsumes many known universal approximators. Building on these, we show that two-layer multi-head attention alone suffices as a sequence-to-sequence universal approximator. In contrast, prior works rely on feed-forward networks to establish universal approximation in Transformers. Furthermore, we extend our techniques to show that, (softmax-)attention-only layers are capable of approximating various statistical models in-context. We believe these techniques hold independent interest.
title Universal Approximation with Softmax Attention
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.15956