Attention as a Hypernetwork
Fuente:
arXiv
Saved in:
| Main Authors: | Schug, Simon, Kobayashi, Seijin, Akram, Yassir, Sacramento, João, Pascanu, Razvan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When can transformers compositionally generalize in-context?
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023)
by: Schug, Simon, et al.
Published: (2023)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Learning Randomized Algorithms with Transformers
by: von Oswald, Johannes, et al.
Published: (2024)
by: von Oswald, Johannes, et al.
Published: (2024)
Scaling can lead to compositional generalization
by: Redhardt, Florian, et al.
Published: (2025)
by: Redhardt, Florian, et al.
Published: (2025)
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023)
by: Zucchet, Nicolas, et al.
Published: (2023)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Uncovering mesa-optimization algorithms in Transformers
by: von Oswald, Johannes, et al.
Published: (2023)
by: von Oswald, Johannes, et al.
Published: (2023)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
What Can Grokking Teach Us About Learning Under Nonstationarity?
by: Lyle, Clare, et al.
Published: (2025)
by: Lyle, Clare, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
by: von Oswald, Johannes, et al.
Published: (2025)
by: von Oswald, Johannes, et al.
Published: (2025)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
by: Moalla, Skander, et al.
Published: (2024)
by: Moalla, Skander, et al.
Published: (2024)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
Hadamard product in deep learning: Introduction, Advances and Challenges
by: Chrysos, Grigorios G, et al.
Published: (2025)
by: Chrysos, Grigorios G, et al.
Published: (2025)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
by: Diep, Nghiem T., et al.
Published: (2025)
by: Diep, Nghiem T., et al.
Published: (2025)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
by: Schmied, Thomas, et al.
Published: (2025)
by: Schmied, Thomas, et al.
Published: (2025)
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
by: Masarczyk, Wojciech, et al.
Published: (2025)
by: Masarczyk, Wojciech, et al.
Published: (2025)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
by: Dufort-Labbé, Simon, et al.
Published: (2024)
by: Dufort-Labbé, Simon, et al.
Published: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)
by: Orvieto, Antonio, et al.
Published: (2023)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
Mining Generalizable Activation Functions
by: Vitvitskyi, Alex, et al.
Published: (2026)
by: Vitvitskyi, Alex, et al.
Published: (2026)
The Illusion of Stochasticity in LLMs
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
by: Lewis, Owen, et al.
Published: (2025)
by: Lewis, Owen, et al.
Published: (2025)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
Disentangling the Causes of Plasticity Loss in Neural Networks
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Hypernetworks for Perspectivist Adaptation
by: Ignatev, Daniil, et al.
Published: (2025)
by: Ignatev, Daniil, et al.
Published: (2025)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
by: Wołczyk, Maciej, et al.
Published: (2024)
by: Wołczyk, Maciej, et al.
Published: (2024)
Transformers need glasses! Information over-squashing in language tasks
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Hypernetworks for Dynamic Feature Selection
by: Fumanal-Idocin, Javier, et al.
Published: (2026)
by: Fumanal-Idocin, Javier, et al.
Published: (2026)
Similar Items
-
When can transformers compositionally generalize in-context?
by: Kobayashi, Seijin, et al.
Published: (2024) -
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023) -
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024) -
Learning Randomized Algorithms with Transformers
by: von Oswald, Johannes, et al.
Published: (2024) -
Scaling can lead to compositional generalization
by: Redhardt, Florian, et al.
Published: (2025)