Saved in:
| Main Authors: | Dragutinović, Sara, Saxe, Andrew M., Singh, Aaditya K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.10425 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
by: Singh, Aaditya K., et al.
Published: (2025)
by: Singh, Aaditya K., et al.
Published: (2025)
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
On the rate of convergence of an over-parametrized Transformer classifier learned by gradient descent
by: Kohler, Michael, et al.
Published: (2023)
by: Kohler, Michael, et al.
Published: (2023)
The broader spectrum of in-context learning
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
by: Lampinen, Andrew Kyle, et al.
Published: (2024)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
by: Dragutinović, Sara, et al.
Published: (2026)
by: Dragutinović, Sara, et al.
Published: (2026)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
State-space models can learn in-context by gradient descent
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
Analysis of the rate of convergence of an over-parametrized convolutional neural network image classifier learned by gradient descent
by: Kohler, Michael, et al.
Published: (2024)
by: Kohler, Michael, et al.
Published: (2024)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Gradient descent for deep equilibrium single-index models
by: Dandapanthula, Sanjit, et al.
Published: (2025)
by: Dandapanthula, Sanjit, et al.
Published: (2025)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Spike-timing-dependent Hebbian learning as noisy gradient descent
by: Dexheimer, Niklas, et al.
Published: (2025)
by: Dexheimer, Niklas, et al.
Published: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Double descent in quantum kernel methods
by: Kempkes, Marie, et al.
Published: (2025)
by: Kempkes, Marie, et al.
Published: (2025)
Early learning of the optimal constant solution in neural networks and humans
by: Rubruck, Jirko, et al.
Published: (2024)
by: Rubruck, Jirko, et al.
Published: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
Shot-frugal and Robust quantum kernel classifiers
by: Shastry, Abhay, et al.
Published: (2022)
by: Shastry, Abhay, et al.
Published: (2022)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Training thermodynamic computers by gradient descent
by: Whitelam, Stephen
Published: (2025)
by: Whitelam, Stephen
Published: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
by: van Rossem, Loek, et al.
Published: (2025)
by: van Rossem, Loek, et al.
Published: (2025)
When Representations Align: Universality in Representation Learning Dynamics
by: van Rossem, Loek, et al.
Published: (2024)
by: van Rossem, Loek, et al.
Published: (2024)
Nonlinear dynamics of localization in neural receptive fields
by: Lufkin, Leon, et al.
Published: (2025)
by: Lufkin, Leon, et al.
Published: (2025)
New logarithmic step size for stochastic gradient descent
by: Shamaee, M. Soheil, et al.
Published: (2024)
by: Shamaee, M. Soheil, et al.
Published: (2024)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
by: Gonsior, Julius, et al.
Published: (2022)
by: Gonsior, Julius, et al.
Published: (2022)
Singular-limit analysis of gradient descent with noise injection
by: Shalova, Anna, et al.
Published: (2024)
by: Shalova, Anna, et al.
Published: (2024)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
by: Beaglehole, Daniel, et al.
Published: (2024)
by: Beaglehole, Daniel, et al.
Published: (2024)
Operationalizing Stein's Method for Online Linear Optimization: CLT-Based Optimal Tradeoffs
by: Zhang, Zhiyu, et al.
Published: (2026)
by: Zhang, Zhiyu, et al.
Published: (2026)
GradCheck: Analyzing classifier guidance gradients for conditional diffusion sampling
by: Vaeth, Philipp, et al.
Published: (2024)
by: Vaeth, Philipp, et al.
Published: (2024)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Generalisation under gradient descent via deterministic PAC-Bayes
by: Clerico, Eugenio, et al.
Published: (2022)
by: Clerico, Eugenio, et al.
Published: (2022)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Softmax Transformers are Turing-Complete
by: Jiang, Hongjian, et al.
Published: (2025)
by: Jiang, Hongjian, et al.
Published: (2025)
Transformer learns the cross-task prior and regularization for in-context learning
by: Lu, Fei, et al.
Published: (2025)
by: Lu, Fei, et al.
Published: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Similar Items
-
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
by: Singh, Aaditya K., et al.
Published: (2025) -
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
by: Singh, Aaditya K., et al.
Published: (2024) -
Training Dynamics of In-Context Learning in Linear Attention
by: Zhang, Yedi, et al.
Published: (2025) -
On the rate of convergence of an over-parametrized Transformer classifier learned by gradient descent
by: Kohler, Michael, et al.
Published: (2023) -
The broader spectrum of in-context learning
by: Lampinen, Andrew Kyle, et al.
Published: (2024)