Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sheen, Heejune, Chen, Siyu, Wang, Tianhao, Zhou, Harrison H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
von: Regis, Eric, et al.
Veröffentlicht: (2026)
von: Regis, Eric, et al.
Veröffentlicht: (2026)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
von: Huang, Yu, et al.
Veröffentlicht: (2026)
von: Huang, Yu, et al.
Veröffentlicht: (2026)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
von: Varre, Aditya, et al.
Veröffentlicht: (2026)
von: Varre, Aditya, et al.
Veröffentlicht: (2026)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
von: Ozaslan, Ibrahim K., et al.
Veröffentlicht: (2024)
von: Ozaslan, Ibrahim K., et al.
Veröffentlicht: (2024)
How Well Can Transformers Emulate In-context Newton's Method?
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
von: Giannou, Angeliki, et al.
Veröffentlicht: (2024)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2025)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
Delightful Policy Gradient
von: Osband, Ian
Veröffentlicht: (2026)
von: Osband, Ian
Veröffentlicht: (2026)
On the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
von: Chen, Lesi, et al.
Veröffentlicht: (2023)
von: Chen, Lesi, et al.
Veröffentlicht: (2023)
Explicit and data-Efficient Encoding via Gradient Flow
von: Flouris, Kyriakos, et al.
Veröffentlicht: (2024)
von: Flouris, Kyriakos, et al.
Veröffentlicht: (2024)
Gradient Descent Efficiency Index
von: Dhingra, Aviral
Veröffentlicht: (2024)
von: Dhingra, Aviral
Veröffentlicht: (2024)
Delightful Distributed Policy Gradient
von: Osband, Ian
Veröffentlicht: (2026)
von: Osband, Ian
Veröffentlicht: (2026)
Robust Control with Gradient Uncertainty
von: Qi, Qian
Veröffentlicht: (2025)
von: Qi, Qian
Veröffentlicht: (2025)
Unbiased Gradient Low-Rank Projection
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
Stability of Transformers under Layer Normalization
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
von: Kan, Kelvin, et al.
Veröffentlicht: (2025)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
Riemann Sum Optimization for Accurate Integrated Gradients Computation
von: Swain, Swadesh, et al.
Veröffentlicht: (2024)
von: Swain, Swadesh, et al.
Veröffentlicht: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
von: Basu, Debabrota, et al.
Veröffentlicht: (2025)
von: Basu, Debabrota, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
von: Arda, Enes, et al.
Veröffentlicht: (2026)
von: Arda, Enes, et al.
Veröffentlicht: (2026)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
von: Hennick, Max, et al.
Veröffentlicht: (2025)
von: Hennick, Max, et al.
Veröffentlicht: (2025)
Tree-Preconditioned Differentiable Optimization and Axioms as Layers
von: Liao, Yuexin
Veröffentlicht: (2025)
von: Liao, Yuexin
Veröffentlicht: (2025)
Nash Equilibria, Regularization and Computation in Optimal Transport-Based Distributionally Robust Optimization
von: Shafiee, Soroosh, et al.
Veröffentlicht: (2023)
von: Shafiee, Soroosh, et al.
Veröffentlicht: (2023)
On the Optimal Construction of Unbiased Gradient Estimators for Zeroth-Order Optimization
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
von: Ma, Shaocong, et al.
Veröffentlicht: (2025)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
von: Hu, Zeou, et al.
Veröffentlicht: (2026)
von: Hu, Zeou, et al.
Veröffentlicht: (2026)
Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
von: Cheng, Ziheng, et al.
Veröffentlicht: (2025)
von: Cheng, Ziheng, et al.
Veröffentlicht: (2025)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2026)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2026)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
von: Ghosh, Ipsita, et al.
Veröffentlicht: (2025)
von: Ghosh, Ipsita, et al.
Veröffentlicht: (2025)
Graph Similarity Regularized Softmax for Semi-Supervised Node Classification
von: Yang, Yiming, et al.
Veröffentlicht: (2024)
von: Yang, Yiming, et al.
Veröffentlicht: (2024)
Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
von: Liu, Xin, et al.
Veröffentlicht: (2022)
von: Liu, Xin, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
von: Chen, Siyu, et al.
Veröffentlicht: (2024) -
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024) -
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
von: Li, Zihao, et al.
Veröffentlicht: (2024) -
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
von: Chen, Siyu, et al.
Veröffentlicht: (2025) -
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
von: Regis, Eric, et al.
Veröffentlicht: (2026)