Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Chen, Schmidt, Mark, Thrampoulidis, Christos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Implicit Bias and Fast Convergence Rates for Self-attention
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
von: Xie, Shengping, et al.
Veröffentlicht: (2026)
von: Xie, Shengping, et al.
Veröffentlicht: (2026)
Implicit Bias of Mirror Flow on Separable Data
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
von: Pesme, Scott, et al.
Veröffentlicht: (2024)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2026)
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2026)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
von: Jung, Hyunji, et al.
Veröffentlicht: (2025)
von: Jung, Hyunji, et al.
Veröffentlicht: (2025)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
von: Garrod, Connall, et al.
Veröffentlicht: (2025)
von: Garrod, Connall, et al.
Veröffentlicht: (2025)
On the Optimization and Generalization of Multi-head Attention
von: Deora, Puneesh, et al.
Veröffentlicht: (2023)
von: Deora, Puneesh, et al.
Veröffentlicht: (2023)
Muon Optimizes Under Spectral Norm Constraints
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Memory capacity of two layer neural networks with smooth activations
von: Madden, Liam, et al.
Veröffentlicht: (2023)
von: Madden, Liam, et al.
Veröffentlicht: (2023)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
Preconditioning Benefits of Spectral Orthogonalization in Muon
von: Ma, Jianhao, et al.
Veröffentlicht: (2026)
von: Ma, Jianhao, et al.
Veröffentlicht: (2026)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
von: Takezawa, Yuki, et al.
Veröffentlicht: (2025)
von: Takezawa, Yuki, et al.
Veröffentlicht: (2025)
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
von: Karnik, Santhosh, et al.
Veröffentlicht: (2024)
von: Karnik, Santhosh, et al.
Veröffentlicht: (2024)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
von: Ravi, Hrithik, et al.
Veröffentlicht: (2024)
von: Ravi, Hrithik, et al.
Veröffentlicht: (2024)
Convergence of Spectral Descent for Non-smooth Optimization
von: Yang, Yixuan, et al.
Veröffentlicht: (2026)
von: Yang, Yixuan, et al.
Veröffentlicht: (2026)
LiMuon: Light and Fast Muon Optimizer for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
von: Huang, Feihu, et al.
Veröffentlicht: (2025)
Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data
von: Tsoy, Nikita, et al.
Veröffentlicht: (2024)
von: Tsoy, Nikita, et al.
Veröffentlicht: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
Muon Dynamics as a Spectral Wasserstein Flow
von: Peyré, Gabriel
Veröffentlicht: (2026)
von: Peyré, Gabriel
Veröffentlicht: (2026)
Next-token prediction capacity: general upper bounds and a lower bound for transformers
von: Madden, Liam, et al.
Veröffentlicht: (2024)
von: Madden, Liam, et al.
Veröffentlicht: (2024)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
von: Li, Tianyou, et al.
Veröffentlicht: (2023)
von: Li, Tianyou, et al.
Veröffentlicht: (2023)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
von: Sato, Naoki, et al.
Veröffentlicht: (2023)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
MuonBP: Faster Muon via Block-Periodic Orthogonalization
von: Khaled, Ahmed, et al.
Veröffentlicht: (2025)
von: Khaled, Ahmed, et al.
Veröffentlicht: (2025)
Towards a Principled Muon under $μ\mathsf{P}$: Ensuring Spectral Conditions throughout Training
von: Zhao, John
Veröffentlicht: (2026)
von: Zhao, John
Veröffentlicht: (2026)
Transformers as Support Vector Machines
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
von: Tarzanagh, Davoud Ataee, et al.
Veröffentlicht: (2023)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
Stochastic Gradient Descent with Adaptive Data
von: Che, Ethan, et al.
Veröffentlicht: (2024)
von: Che, Ethan, et al.
Veröffentlicht: (2024)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
von: Hou, Yikun, et al.
Veröffentlicht: (2025)
von: Hou, Yikun, et al.
Veröffentlicht: (2025)
Implicit Regularization in Perturbed Deep Matrix Factorization: Spectral Conditions and Stability
von: Wang, Jingzhe, et al.
Veröffentlicht: (2026)
von: Wang, Jingzhe, et al.
Veröffentlicht: (2026)
On the Implicit Bias of Adam
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
von: Cattaneo, Matias D., et al.
Veröffentlicht: (2023)
Error Feedback for Muon and Friends
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
von: Gruntkowska, Kaja, et al.
Veröffentlicht: (2025)
Convergence of Muon with Newton-Schulz
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
von: Kim, Gyu Yeol, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Implicit Bias and Fast Convergence Rates for Self-attention
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2024) -
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
von: Xie, Shengping, et al.
Veröffentlicht: (2026) -
Implicit Bias of Mirror Flow on Separable Data
von: Pesme, Scott, et al.
Veröffentlicht: (2024) -
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
von: Akhtiamov, Danil, et al.
Veröffentlicht: (2026) -
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)