The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Otsuka, Hikari, Chijiwa, Daiki, Okoshi, Yasuyuki, Fujiki, Daichi, Takeuchi, Susumu, Motomura, Masato
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911251276234752
author Otsuka, Hikari
Chijiwa, Daiki
Okoshi, Yasuyuki
Fujiki, Daichi
Takeuchi, Susumu
Motomura, Masato
author_facet Otsuka, Hikari
Chijiwa, Daiki
Okoshi, Yasuyuki
Fujiki, Daichi
Takeuchi, Susumu
Motomura, Masato
contents The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks. Although recent theoretical studies have established the SLTH across various neural architectures, the SLTH for transformer architectures still lacks theoretical understanding. In particular, the current theory of the SLTH does not yet account for the multi-head attention (MHA) mechanism, a core component of transformers. To address this gap, we introduce a theoretical analysis of the existence of SLTs within MHAs. We prove that, if a randomly initialized MHA of $H$ heads and input dimension $d$ has the hidden dimension $O(d\log(Hd^{3/2}))$ for the key and value, it contains an SLT that approximates an arbitrary MHA with the same input dimension with high probability. Furthermore, by leveraging this theory for MHAs, we extend the SLTH to transformers without normalization layers. We empirically validate our theoretical findings, demonstrating that the approximation error between the SLT within a source model (MHA and transformer) and an approximate target counterpart decreases exponentially by increasing the hidden dimension of the source model.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
Otsuka, Hikari
Chijiwa, Daiki
Okoshi, Yasuyuki
Fujiki, Daichi
Takeuchi, Susumu
Motomura, Masato
Machine Learning
Artificial Intelligence
The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks. Although recent theoretical studies have established the SLTH across various neural architectures, the SLTH for transformer architectures still lacks theoretical understanding. In particular, the current theory of the SLTH does not yet account for the multi-head attention (MHA) mechanism, a core component of transformers. To address this gap, we introduce a theoretical analysis of the existence of SLTs within MHAs. We prove that, if a randomly initialized MHA of $H$ heads and input dimension $d$ has the hidden dimension $O(d\log(Hd^{3/2}))$ for the key and value, it contains an SLT that approximates an arbitrary MHA with the same input dimension with high probability. Furthermore, by leveraging this theory for MHAs, we extend the SLTH to transformers without normalization layers. We empirically validate our theoretical findings, demonstrating that the approximation error between the SLT within a source model (MHA and transformer) and an approximate target counterpart decreases exponentially by increasing the hidden dimension of the source model.
title The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.04217