Training Dynamics of In-Context Learning in Linear Attention
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yedi, Singh, Aaditya K., Latham, Peter E., Saxe, Andrew |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Understanding Unimodal Bias in Multimodal Deep Linear Networks
par: Zhang, Yedi, et autres
Publié: (2023)
par: Zhang, Yedi, et autres
Publié: (2023)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
par: Zhang, Yedi, et autres
Publié: (2024)
par: Zhang, Yedi, et autres
Publié: (2024)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
par: Zhang, Yedi, et autres
Publié: (2025)
par: Zhang, Yedi, et autres
Publié: (2025)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
par: Dragutinović, Sara, et autres
Publié: (2025)
par: Dragutinović, Sara, et autres
Publié: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
par: Lee, Jin Hwa, et autres
Publié: (2025)
par: Lee, Jin Hwa, et autres
Publié: (2025)
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
par: Singh, Aaditya K., et autres
Publié: (2025)
par: Singh, Aaditya K., et autres
Publié: (2025)
Optimal Learning Rate Schedule for Balancing Effort and Performance
par: Njaradi, Valentina, et autres
Publié: (2026)
par: Njaradi, Valentina, et autres
Publié: (2026)
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
par: Singh, Aaditya K., et autres
Publié: (2024)
par: Singh, Aaditya K., et autres
Publié: (2024)
When Representations Align: Universality in Representation Learning Dynamics
par: van Rossem, Loek, et autres
Publié: (2024)
par: van Rossem, Loek, et autres
Publié: (2024)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
par: Joshi, Sahil, et autres
Publié: (2025)
par: Joshi, Sahil, et autres
Publié: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
par: He, Jianliang, et autres
Publié: (2025)
par: He, Jianliang, et autres
Publié: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
par: Dominé, Clémentine C. J., et autres
Publié: (2024)
par: Dominé, Clémentine C. J., et autres
Publié: (2024)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
par: Pandey, Vishal, et autres
Publié: (2026)
par: Pandey, Vishal, et autres
Publié: (2026)
Operationalizing Stein's Method for Online Linear Optimization: CLT-Based Optimal Tradeoffs
par: Zhang, Zhiyu, et autres
Publié: (2026)
par: Zhang, Zhiyu, et autres
Publié: (2026)
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers
par: Chen, Brian K, et autres
Publié: (2024)
par: Chen, Brian K, et autres
Publié: (2024)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
par: Njaradi, Valentina, et autres
Publié: (2026)
par: Njaradi, Valentina, et autres
Publié: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
par: Xie, Zixuan, et autres
Publié: (2026)
par: Xie, Zixuan, et autres
Publié: (2026)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
par: Singh, Aaditya K., et autres
Publié: (2024)
par: Singh, Aaditya K., et autres
Publié: (2024)
InAttention: Linear Context Scaling for Transformers
par: Eisner, Joseph
Publié: (2024)
par: Eisner, Joseph
Publié: (2024)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
par: Patel, Nishil, et autres
Publié: (2023)
par: Patel, Nishil, et autres
Publié: (2023)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
par: Jarvis, Devon, et autres
Publié: (2025)
par: Jarvis, Devon, et autres
Publié: (2025)
Nonlinear dynamics of localization in neural receptive fields
par: Lufkin, Leon, et autres
Publié: (2025)
par: Lufkin, Leon, et autres
Publié: (2025)
In-Context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-Separation
par: Cole, Frank, et autres
Publié: (2025)
par: Cole, Frank, et autres
Publié: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
par: van Rossem, Loek, et autres
Publié: (2025)
par: van Rossem, Loek, et autres
Publié: (2025)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
par: Xu, Hongtao, et autres
Publié: (2026)
par: Xu, Hongtao, et autres
Publié: (2026)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
par: Jiang, Jiarui, et autres
Publié: (2025)
par: Jiang, Jiarui, et autres
Publié: (2025)
Logarithmic Neyman Regret for Adaptive Estimation of the Average Treatment Effect
par: Neopane, Ojash, et autres
Publié: (2024)
par: Neopane, Ojash, et autres
Publié: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
par: Cui, Yingqian, et autres
Publié: (2024)
par: Cui, Yingqian, et autres
Publié: (2024)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
par: Goel, Gautam, et autres
Publié: (2026)
par: Goel, Gautam, et autres
Publié: (2026)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
par: Chen, Siyu, et autres
Publié: (2024)
par: Chen, Siyu, et autres
Publié: (2024)
Meta-Learning Strategies through Value Maximization in Neural Networks
par: Carrasco-Davis, Rodrigo, et autres
Publié: (2023)
par: Carrasco-Davis, Rodrigo, et autres
Publié: (2023)
In-context Learning for Mixture of Linear Regressions: Existence, Generalization and Training Dynamics
par: Jin, Yanhao, et autres
Publié: (2024)
par: Jin, Yanhao, et autres
Publié: (2024)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
par: Goel, Ayush, et autres
Publié: (2026)
par: Goel, Ayush, et autres
Publié: (2026)
Early learning of the optimal constant solution in neural networks and humans
par: Rubruck, Jirko, et autres
Publié: (2024)
par: Rubruck, Jirko, et autres
Publié: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
par: Li, Wenxuan, et autres
Publié: (2025)
par: Li, Wenxuan, et autres
Publié: (2025)
State Rank Dynamics in Linear Attention LLMs
par: Sun, Ao, et autres
Publié: (2026)
par: Sun, Ao, et autres
Publié: (2026)
Optimistic Algorithms for Adaptive Estimation of the Average Treatment Effect
par: Neopane, Ojash, et autres
Publié: (2025)
par: Neopane, Ojash, et autres
Publié: (2025)
Agents Learn Their Runtime: Interpreter Persistence as Training-Time Semantics
par: May, Victor, et autres
Publié: (2026)
par: May, Victor, et autres
Publié: (2026)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
par: Chen, Xingwu, et autres
Publié: (2024)
par: Chen, Xingwu, et autres
Publié: (2024)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
par: Kim, Juno, et autres
Publié: (2024)
par: Kim, Juno, et autres
Publié: (2024)
Documents similaires
-
Understanding Unimodal Bias in Multimodal Deep Linear Networks
par: Zhang, Yedi, et autres
Publié: (2023) -
When Are Bias-Free ReLU Networks Effectively Linear Networks?
par: Zhang, Yedi, et autres
Publié: (2024) -
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
par: Zhang, Yedi, et autres
Publié: (2025) -
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
par: Dragutinović, Sara, et autres
Publié: (2025) -
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
par: Lee, Jin Hwa, et autres
Publié: (2025)