Implicit Bias of Mirror Flow on Separable Data
Fuente:
arXiv
Saved in:
| Main Authors: | Pesme, Scott, Dragomir, Radu-Alexandru, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026)
by: Varre, Aditya, et al.
Published: (2026)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
MAP Estimation with Denoisers: Convergence Rates and Guarantees
by: Pesme, Scott, et al.
Published: (2025)
by: Pesme, Scott, et al.
Published: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data
by: Tsoy, Nikita, et al.
Published: (2024)
by: Tsoy, Nikita, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
Importance Sampling Optimization with Laplace Principle
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
by: Dragomir, Radu-Alexandru, et al.
Published: (2026)
Consensus-Based Optimization Beyond Finite-Time Analysis
by: Bianchi, Pascal, et al.
Published: (2025)
by: Bianchi, Pascal, et al.
Published: (2025)
Convex quartic problems: homogenized gradient method and preconditioning
by: Dragomir, Radu-Alexandru, et al.
Published: (2023)
by: Dragomir, Radu-Alexandru, et al.
Published: (2023)
Implicit Bias in Matrix Factorization and its Explicit Realization in a New Architecture
by: Hou, Yikun, et al.
Published: (2025)
by: Hou, Yikun, et al.
Published: (2025)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
The Implicit Bias of Heterogeneity towards Invariance: A Study of Multi-Environment Matrix Sensing
by: Xu, Yang, et al.
Published: (2024)
by: Xu, Yang, et al.
Published: (2024)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
Computing the Bias of Constant-step Stochastic Approximation with Markovian Noise
by: Allmeier, Sebastian, et al.
Published: (2024)
by: Allmeier, Sebastian, et al.
Published: (2024)
Parameter-free Mirror Descent
by: Jacobsen, Andrew, et al.
Published: (2022)
by: Jacobsen, Andrew, et al.
Published: (2022)
Mirror Descent on Riemannian Manifolds
by: Jiang, Jiaxin, et al.
Published: (2026)
by: Jiang, Jiaxin, et al.
Published: (2026)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
by: Lai, Kuo-Wei, et al.
Published: (2026)
by: Lai, Kuo-Wei, et al.
Published: (2026)
Policy Mirror Descent with Temporal Difference Learning: Sample Complexity under Online Markov Data
by: Li, Wenye, et al.
Published: (2025)
by: Li, Wenye, et al.
Published: (2025)
Mirror Mean-Field Langevin Dynamics
by: Gu, Anming, et al.
Published: (2025)
by: Gu, Anming, et al.
Published: (2025)
Mirror Descent on Reproducing Kernel Banach Spaces
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
Mirror and Preconditioned Gradient Descent in Wasserstein Space
by: Bonet, Clément, et al.
Published: (2024)
by: Bonet, Clément, et al.
Published: (2024)
On the Convergence of Policy in Unregularized Policy Mirror Descent
by: Lin, Dachao, et al.
Published: (2022)
by: Lin, Dachao, et al.
Published: (2022)
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025)
by: Crăciun, Alexandru, et al.
Published: (2025)
The Effect of Mini-Batch Noise on the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2026)
by: Cattaneo, Matias D., et al.
Published: (2026)
A Mirror Descent Perspective of Smoothed Sign Descent
by: Wang, Shuyang, et al.
Published: (2024)
by: Wang, Shuyang, et al.
Published: (2024)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
by: Liu, Jiacai, et al.
Published: (2025)
by: Liu, Jiacai, et al.
Published: (2025)
Optimizer's Information Criterion: Dissecting and Correcting Bias in Data-Driven Optimization
by: Iyengar, Garud, et al.
Published: (2023)
by: Iyengar, Garud, et al.
Published: (2023)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Symplectic Inductive Bias for Data-Driven Target Reachability in Hamiltonian Systems
by: Ouyang, Zhuo, et al.
Published: (2026)
by: Ouyang, Zhuo, et al.
Published: (2026)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Similar Items
-
A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation
by: Boursier, Etienne, et al.
Published: (2025) -
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024) -
Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
by: Varre, Aditya, et al.
Published: (2026) -
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026) -
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)