Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fasina, Oluwadamilola, Pohle, Ruben V. C., Su, Pei-Chun, Coifman, Ronald R.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908412819800064
author Fasina, Oluwadamilola
Pohle, Ruben V. C.
Su, Pei-Chun
Coifman, Ronald R.
author_facet Fasina, Oluwadamilola
Pohle, Ruben V. C.
Su, Pei-Chun
Coifman, Ronald R.
contents We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechanism to softmax activation is obtained by appealing to paradifferential calculus, (and is supported by computational examples), which relies on the intrinsic organization of the attention heads. Furthermore, we use an existing methodology for hierarchical organization of tensors to examine network structure by constructing hierarchal partition trees with respect to the query, key, and head axes of network 3-tensors. Such an organization is consequential since it allows one to profitably execute common signal processing tasks on a geometry where the organized network 3-tensors exhibit regularity. We exemplify this qualitatively, by visualizing the hierarchical organization of the tree comprised of attention heads and the diffusion map embeddings, and quantitatively by investigating network sparsity with the expansion coefficients of individual attention heads and the entire network with respect to the bi and tri-haar bases (respectively) on the space of queries, keys, and heads of the network. To showcase the utility of our theoretical and methodological findings, we provide computational examples using vision and language transformers. The ramifications of these findings are two-fold: (1) a subsequent step in interpretability analysis is theoretically admitted, and can be exploited empirically for downstream interpretability tasks (2) one can use the network 3-tensor organization for empirical network applications such as model pruning (by virtue of network sparsity) and network architecture comparison.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity
Fasina, Oluwadamilola
Pohle, Ruben V. C.
Su, Pei-Chun
Coifman, Ronald R.
Numerical Analysis
Artificial Intelligence
We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechanism to softmax activation is obtained by appealing to paradifferential calculus, (and is supported by computational examples), which relies on the intrinsic organization of the attention heads. Furthermore, we use an existing methodology for hierarchical organization of tensors to examine network structure by constructing hierarchal partition trees with respect to the query, key, and head axes of network 3-tensors. Such an organization is consequential since it allows one to profitably execute common signal processing tasks on a geometry where the organized network 3-tensors exhibit regularity. We exemplify this qualitatively, by visualizing the hierarchical organization of the tree comprised of attention heads and the diffusion map embeddings, and quantitatively by investigating network sparsity with the expansion coefficients of individual attention heads and the entire network with respect to the bi and tri-haar bases (respectively) on the space of queries, keys, and heads of the network. To showcase the utility of our theoretical and methodological findings, we provide computational examples using vision and language transformers. The ramifications of these findings are two-fold: (1) a subsequent step in interpretability analysis is theoretically admitted, and can be exploited empirically for downstream interpretability tasks (2) one can use the network 3-tensor organization for empirical network applications such as model pruning (by virtue of network sparsity) and network architecture comparison.
title Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity
topic Numerical Analysis
Artificial Intelligence
url https://arxiv.org/abs/2506.15541