Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Siyu, Sheen, Heejune, Wang, Tianhao, Yang, Zhuoran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
por: Chen, Siyu, et al.
Publicado: (2024)
por: Chen, Siyu, et al.
Publicado: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024)
por: Sheen, Heejune, et al.
Publicado: (2024)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
por: Chen, Siyu, et al.
Publicado: (2025)
por: Chen, Siyu, et al.
Publicado: (2025)
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
por: Yang, Zhuoran, et al.
Publicado: (2020)
por: Yang, Zhuoran, et al.
Publicado: (2020)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
por: Bareilles, Gilles, et al.
Publicado: (2026)
por: Bareilles, Gilles, et al.
Publicado: (2026)
Statistical and Algorithmic Foundations of Reinforcement Learning
por: Chi, Yuejie, et al.
Publicado: (2025)
por: Chi, Yuejie, et al.
Publicado: (2025)
Sinkhorn Based Associative Memory Retrieval Using Spherical Hellinger Kantorovich Dynamics
por: Mustafi, Aratrika, et al.
Publicado: (2026)
por: Mustafi, Aratrika, et al.
Publicado: (2026)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
por: Rashidinejad, Paria, et al.
Publicado: (2024)
por: Rashidinejad, Paria, et al.
Publicado: (2024)
A Differential and Pointwise Control Approach to Reinforcement Learning
por: Nguyen, Minh, et al.
Publicado: (2024)
por: Nguyen, Minh, et al.
Publicado: (2024)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
por: Nguyen, Minh
Publicado: (2026)
por: Nguyen, Minh
Publicado: (2026)
Inverse Mixed-Integer Programming: Learning Constraints then Objective Functions
por: Kitaoka, Akira
Publicado: (2025)
por: Kitaoka, Akira
Publicado: (2025)
Learning to Fuse Temporal Proximity Networks: A Case Study in Chimpanzee Social Interactions
por: He, Yixuan, et al.
Publicado: (2025)
por: He, Yixuan, et al.
Publicado: (2025)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
por: Mohamed, Mimoun, et al.
Publicado: (2023)
por: Mohamed, Mimoun, et al.
Publicado: (2023)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
por: Han, Qiyang, et al.
Publicado: (2025)
por: Han, Qiyang, et al.
Publicado: (2025)
FraPPE: Fast and Efficient Preference-based Pure Exploration
por: Das, Udvas, et al.
Publicado: (2025)
por: Das, Udvas, et al.
Publicado: (2025)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
por: Yan, Shunxing, et al.
Publicado: (2026)
por: Yan, Shunxing, et al.
Publicado: (2026)
Smooth Non-Stationary Bandits
por: Jia, Su, et al.
Publicado: (2023)
por: Jia, Su, et al.
Publicado: (2023)
Piecewise Polynomial Regression of Tame Functions via Integer Programming
por: Bareilles, Gilles, et al.
Publicado: (2023)
por: Bareilles, Gilles, et al.
Publicado: (2023)
Blessings and Curses of Covariate Shifts: Adversarial Learning Dynamics, Directional Convergence, and Equilibria
por: Liang, Tengyuan
Publicado: (2022)
por: Liang, Tengyuan
Publicado: (2022)
On the Uniform Convergence of Subdifferentials in Stochastic Optimization and Learning
por: Ruan, Feng
Publicado: (2024)
por: Ruan, Feng
Publicado: (2024)
A Graphical Global Optimization Framework for Parameter Estimation of Statistical Models with Nonconvex Regularization Functions
por: Davarnia, Danial, et al.
Publicado: (2025)
por: Davarnia, Danial, et al.
Publicado: (2025)
Denoising Diffusions with Optimal Transport: Localization, Curvature, and Multi-Scale Complexity
por: Liang, Tengyuan, et al.
Publicado: (2024)
por: Liang, Tengyuan, et al.
Publicado: (2024)
Learning an Optimal Assortment Policy under Observational Data
por: Han, Yuxuan, et al.
Publicado: (2025)
por: Han, Yuxuan, et al.
Publicado: (2025)
Learning and Decision-Making with Data: Optimal Formulations and Phase Transitions
por: Bennouna, Amine, et al.
Publicado: (2021)
por: Bennouna, Amine, et al.
Publicado: (2021)
Is Q-Learning Minimax Optimal? A Tight Sample Complexity Analysis
por: Li, Gen, et al.
Publicado: (2021)
por: Li, Gen, et al.
Publicado: (2021)
Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence
por: Chaudhry, Faris, et al.
Publicado: (2026)
por: Chaudhry, Faris, et al.
Publicado: (2026)
Progressive Feedforward Collapse of ResNet Training
por: Wang, Sicong, et al.
Publicado: (2024)
por: Wang, Sicong, et al.
Publicado: (2024)
A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
por: Alfano, Carlo, et al.
Publicado: (2023)
por: Alfano, Carlo, et al.
Publicado: (2023)
A Theory of Feature Learning in Kernel Models
por: Chen, Yunlu, et al.
Publicado: (2023)
por: Chen, Yunlu, et al.
Publicado: (2023)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
por: Cheng, Xiuyuan, et al.
Publicado: (2023)
por: Cheng, Xiuyuan, et al.
Publicado: (2023)
Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees
por: Maros, Marie, et al.
Publicado: (2022)
por: Maros, Marie, et al.
Publicado: (2022)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
High-probability Convergence Bounds for Nonlinear Stochastic Gradient Descent Under Heavy-tailed Noise
por: Armacki, Aleksandar, et al.
Publicado: (2023)
por: Armacki, Aleksandar, et al.
Publicado: (2023)
Conformal Prediction in The Loop: A Feedback-Based Uncertainty Model for Trajectory Optimization
por: Wang, Han, et al.
Publicado: (2025)
por: Wang, Han, et al.
Publicado: (2025)
Stochastic Optimization with Optimal Importance Sampling
por: Aolaritei, Liviu, et al.
Publicado: (2025)
por: Aolaritei, Liviu, et al.
Publicado: (2025)
Accelerating Convergence of Score-Based Diffusion Models, Provably
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
por: Bona-Pellissier, Joachim, et al.
Publicado: (2024)
por: Bona-Pellissier, Joachim, et al.
Publicado: (2024)
Error Analysis of Triangular Optimal Transport Maps for Filtering
por: Al-Jarrah, Mohammad, et al.
Publicado: (2025)
por: Al-Jarrah, Mohammad, et al.
Publicado: (2025)
A New Perspective On Denoising Based On Optimal Transport
por: Trillos, Nicolas Garcia, et al.
Publicado: (2023)
por: Trillos, Nicolas Garcia, et al.
Publicado: (2023)
Optimal transport natural gradient for statistical manifolds with continuous sample space
por: Chen, Yifan, et al.
Publicado: (2018)
por: Chen, Yifan, et al.
Publicado: (2018)
Ejemplares similares
-
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
por: Chen, Siyu, et al.
Publicado: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
por: Sheen, Heejune, et al.
Publicado: (2024) -
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
por: Chen, Siyu, et al.
Publicado: (2025) -
Variational Transport: A Convergent Particle-BasedAlgorithm for Distributional Optimization
por: Yang, Zhuoran, et al.
Publicado: (2020) -
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
por: Bareilles, Gilles, et al.
Publicado: (2026)