Weight decay induces low-rank attention layers
Fuente:
arXiv
Saved in:
| Main Authors: | Kobayashi, Seijin, Akram, Yassir, Von Oswald, Johannes |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Randomized Algorithms with Transformers
by: von Oswald, Johannes, et al.
Published: (2024)
by: von Oswald, Johannes, et al.
Published: (2024)
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023)
by: Zucchet, Nicolas, et al.
Published: (2023)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
When can transformers compositionally generalize in-context?
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023)
by: Schug, Simon, et al.
Published: (2023)
Sinkhorn doubly stochastic attention rank decay analysis
by: Lapenna, Michela, et al.
Published: (2026)
by: Lapenna, Michela, et al.
Published: (2026)
Scaling can lead to compositional generalization
by: Redhardt, Florian, et al.
Published: (2025)
by: Redhardt, Florian, et al.
Published: (2025)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Myosotis: structured computation for attention like layer
by: Egorov, Evgenii, et al.
Published: (2025)
by: Egorov, Evgenii, et al.
Published: (2025)
Uncovering mesa-optimization algorithms in Transformers
by: von Oswald, Johannes, et al.
Published: (2023)
by: von Oswald, Johannes, et al.
Published: (2023)
When Astronomy Meets AI: Manazel For Crescent Visibility Prediction in Morocco
by: Lairgi, Yassir
Published: (2025)
by: Lairgi, Yassir
Published: (2025)
$k$-SVD with Gradient Descent
by: Jedra, Yassir, et al.
Published: (2025)
by: Jedra, Yassir, et al.
Published: (2025)
ARMAX identification of low rank graphical models
by: Cao, Wenqi, et al.
Published: (2025)
by: Cao, Wenqi, et al.
Published: (2025)
SAFE setup for generative molecular design
by: Mesbahi, Yassir El, et al.
Published: (2024)
by: Mesbahi, Yassir El, et al.
Published: (2024)
Provably learning a multi-head attention layer
by: Chen, Sitan, et al.
Published: (2024)
by: Chen, Sitan, et al.
Published: (2024)
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
by: Kobayashi, Seijin, et al.
Published: (2025)
by: Kobayashi, Seijin, et al.
Published: (2025)
Model-free Low-Rank Reinforcement Learning via Leveraged Entry-wise Matrix Estimation
by: Stojanovic, Stefan, et al.
Published: (2024)
by: Stojanovic, Stefan, et al.
Published: (2024)
Low-rank bias, weight decay, and model merging in neural networks
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
Linear Transformers are Versatile In-Context Learners
by: Vladymyrov, Max, et al.
Published: (2024)
by: Vladymyrov, Max, et al.
Published: (2024)
Activation Functions Considered Harmful: Recovering Neural Network Weights through Controlled Channels
by: Spielman, Jesse, et al.
Published: (2025)
by: Spielman, Jesse, et al.
Published: (2025)
Optimal low-rank stochastic gradient estimation for LLM training
by: Li, Zehao, et al.
Published: (2026)
by: Li, Zehao, et al.
Published: (2026)
Worst-case low-rank approximations
by: Fries, Anya, et al.
Published: (2026)
by: Fries, Anya, et al.
Published: (2026)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
A new perspective on low-rank optimization
by: Bertsimas, Dimitris, et al.
Published: (2021)
by: Bertsimas, Dimitris, et al.
Published: (2021)
Efficient Sample-Specific Encoder Perturbations
by: Fathullah, Yassir, et al.
Published: (2024)
by: Fathullah, Yassir, et al.
Published: (2024)
Near-optimal Rank Adaptive Inference of High Dimensional Matrices
by: Zheng, Frédéric, et al.
Published: (2025)
by: Zheng, Frédéric, et al.
Published: (2025)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
by: Fathullah, Yassir, et al.
Published: (2025)
by: Fathullah, Yassir, et al.
Published: (2025)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
by: von Oswald, Johannes, et al.
Published: (2025)
by: von Oswald, Johannes, et al.
Published: (2025)
Expressivity of congruence-based architectures for DNNs on positive-definite matrices
by: Oswald, Antonin, et al.
Published: (2026)
by: Oswald, Antonin, et al.
Published: (2026)
Linear Recursive Feature Machines provably recover low-rank matrices
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
by: Radhakrishnan, Adityanarayanan, et al.
Published: (2024)
A geometric framework for momentum-based optimizers for low-rank training
by: Schotthöfer, Steffen, et al.
Published: (2025)
by: Schotthöfer, Steffen, et al.
Published: (2025)
Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery
by: Jedra, Yassir, et al.
Published: (2024)
by: Jedra, Yassir, et al.
Published: (2024)
Exploiting Observation Bias to Improve Matrix Completion
by: Jedra, Yassir, et al.
Published: (2023)
by: Jedra, Yassir, et al.
Published: (2023)
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
by: Chen, Ke, et al.
Published: (2024)
by: Chen, Ke, et al.
Published: (2024)
Fuel Consumption Prediction: A Comparative Analysis of Machine Learning Paradigms
by: Akram, Ali
Published: (2026)
by: Akram, Ali
Published: (2026)
Wasserstein Gradient Flows for Scalable and Regularized Barycenter Computation
by: Montesuma, Eduardo Fernandes, et al.
Published: (2025)
by: Montesuma, Eduardo Fernandes, et al.
Published: (2025)
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Bitformer: An efficient Transformer with bitwise operation-based attention for Big Data Analytics at low-cost low-precision devices
by: Duan, Gaoxiang, et al.
Published: (2023)
by: Duan, Gaoxiang, et al.
Published: (2023)
Gradient dynamics for low-rank fine-tuning beyond kernels
by: Dayi, Arif Kerem, et al.
Published: (2024)
by: Dayi, Arif Kerem, et al.
Published: (2024)
SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Similar Items
-
Learning Randomized Algorithms with Transformers
by: von Oswald, Johannes, et al.
Published: (2024) -
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023) -
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024) -
When can transformers compositionally generalize in-context?
by: Kobayashi, Seijin, et al.
Published: (2024) -
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023)