Training Dynamics of In-Context Learning in Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yedi, Singh, Aaditya K., Latham, Peter E., Saxe, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023)
by: Zhang, Yedi, et al.
Published: (2023)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025)
by: Dragutinović, Sara, et al.
Published: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)
by: Lee, Jin Hwa, et al.
Published: (2025)
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
by: Singh, Aaditya K., et al.
Published: (2025)
by: Singh, Aaditya K., et al.
Published: (2025)
Optimal Learning Rate Schedule for Balancing Effort and Performance
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
When Representations Align: Universality in Representation Learning Dynamics
by: van Rossem, Loek, et al.
Published: (2024)
by: van Rossem, Loek, et al.
Published: (2024)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
by: He, Jianliang, et al.
Published: (2025)
by: He, Jianliang, et al.
Published: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
by: Dominé, Clémentine C. J., et al.
Published: (2024)
by: Dominé, Clémentine C. J., et al.
Published: (2024)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
by: Pandey, Vishal, et al.
Published: (2026)
by: Pandey, Vishal, et al.
Published: (2026)
Operationalizing Stein's Method for Online Linear Optimization: CLT-Based Optimal Tradeoffs
by: Zhang, Zhiyu, et al.
Published: (2026)
by: Zhang, Zhiyu, et al.
Published: (2026)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
by: Chen, Brian K, et al.
Published: (2024)
by: Chen, Brian K, et al.
Published: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
by: Njaradi, Valentina, et al.
Published: (2026)
by: Njaradi, Valentina, et al.
Published: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
InAttention: Linear Context Scaling for Transformers
by: Eisner, Joseph
Published: (2024)
by: Eisner, Joseph
Published: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023)
by: Patel, Nishil, et al.
Published: (2023)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
by: Jarvis, Devon, et al.
Published: (2025)
by: Jarvis, Devon, et al.
Published: (2025)
Nonlinear dynamics of localization in neural receptive fields
by: Lufkin, Leon, et al.
Published: (2025)
by: Lufkin, Leon, et al.
Published: (2025)
In-Context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-Separation
by: Cole, Frank, et al.
Published: (2025)
by: Cole, Frank, et al.
Published: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
by: van Rossem, Loek, et al.
Published: (2025)
by: van Rossem, Loek, et al.
Published: (2025)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
Logarithmic Neyman Regret for Adaptive Estimation of the Average Treatment Effect
by: Neopane, Ojash, et al.
Published: (2024)
by: Neopane, Ojash, et al.
Published: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
by: Goel, Gautam, et al.
Published: (2026)
by: Goel, Gautam, et al.
Published: (2026)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Meta-Learning Strategies through Value Maximization in Neural Networks
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
by: Carrasco-Davis, Rodrigo, et al.
Published: (2023)
In-context Learning for Mixture of Linear Regressions: Existence, Generalization and Training Dynamics
by: Jin, Yanhao, et al.
Published: (2024)
by: Jin, Yanhao, et al.
Published: (2024)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
by: Goel, Ayush, et al.
Published: (2026)
by: Goel, Ayush, et al.
Published: (2026)
Early learning of the optimal constant solution in neural networks and humans
by: Rubruck, Jirko, et al.
Published: (2024)
by: Rubruck, Jirko, et al.
Published: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
State Rank Dynamics in Linear Attention LLMs
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Optimistic Algorithms for Adaptive Estimation of the Average Treatment Effect
by: Neopane, Ojash, et al.
Published: (2025)
by: Neopane, Ojash, et al.
Published: (2025)
Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics
by: May, Victor, et al.
Published: (2026)
by: May, Victor, et al.
Published: (2026)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Similar Items
-
Understanding Unimodal Bias in Multimodal Deep Linear Networks
by: Zhang, Yedi, et al.
Published: (2023) -
When Are Bias-Free ReLU Networks Effectively Linear Networks?
by: Zhang, Yedi, et al.
Published: (2024) -
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025) -
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
by: Dragutinović, Sara, et al.
Published: (2025) -
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
by: Lee, Jin Hwa, et al.
Published: (2025)