A unified framework for establishing the universal approximation of transformer-type architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Jingpu, Lin, Ting, Shen, Zuowei, Li, Qianxiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917030026805248
author Cheng, Jingpu
Lin, Ting
Shen, Zuowei
Li, Qianxiao
author_facet Cheng, Jingpu
Lin, Ting
Shen, Zuowei
Li, Qianxiao
contents We investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requirement for UAP and introduces a general sufficient condition that applies to a broad class of architectures. Leveraging an analyticity assumption on the attention layer, we can significantly simplify the verification of this condition, providing a non-constructive approach in establishing UAP for such architectures. We demonstrate the applicability of our framework by proving UAP for transformers with various attention mechanisms, including kernel-based and sparse attention mechanisms. The corollaries of our results either generalize prior works or establish UAP for architectures not previously covered. Furthermore, our framework offers a principled foundation for designing novel transformer architectures with inherent UAP guarantees, including those with specific functional symmetries. We propose examples to illustrate these insights.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A unified framework for establishing the universal approximation of transformer-type architectures
Cheng, Jingpu
Lin, Ting
Shen, Zuowei
Li, Qianxiao
Machine Learning
Optimization and Control
We investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requirement for UAP and introduces a general sufficient condition that applies to a broad class of architectures. Leveraging an analyticity assumption on the attention layer, we can significantly simplify the verification of this condition, providing a non-constructive approach in establishing UAP for such architectures. We demonstrate the applicability of our framework by proving UAP for transformers with various attention mechanisms, including kernel-based and sparse attention mechanisms. The corollaries of our results either generalize prior works or establish UAP for architectures not previously covered. Furthermore, our framework offers a principled foundation for designing novel transformer architectures with inherent UAP guarantees, including those with specific functional symmetries. We propose examples to illustrate these insights.
title A unified framework for establishing the universal approximation of transformer-type architectures
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2506.23551