Gradient Flow Polarizes Softmax Outputs towards Low-Entropy Solutions
Fuente:
arXiv
Saved in:
| Main Authors: | Varre, Aditya, Rofin, Mark, Flammarion, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026)
by: Rofin, Mark, et al.
Published: (2026)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods
by: Klein, Sara, et al.
Published: (2023)
by: Klein, Sara, et al.
Published: (2023)
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
by: Varre, Aditya, et al.
Published: (2025)
by: Varre, Aditya, et al.
Published: (2025)
Model-Free Output Feedback Stabilization via Policy Gradient Methods
by: Zhang, Ankang, et al.
Published: (2026)
by: Zhang, Ankang, et al.
Published: (2026)
Graph Similarity Regularized Softmax for Semi-Supervised Node Classification
by: Yang, Yiming, et al.
Published: (2024)
by: Yang, Yiming, et al.
Published: (2024)
GANs as Gradient Flows that Converge
by: Huang, Yu-Jui, et al.
Published: (2022)
by: Huang, Yu-Jui, et al.
Published: (2022)
Flowing Datasets with Wasserstein over Wasserstein Gradient Flows
by: Bonet, Clément, et al.
Published: (2025)
by: Bonet, Clément, et al.
Published: (2025)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
Safe Gradient Flow for Bilevel Optimization
by: Sharifi, Sina, et al.
Published: (2025)
by: Sharifi, Sina, et al.
Published: (2025)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
by: Cayci, Semih, et al.
Published: (2021)
by: Cayci, Semih, et al.
Published: (2021)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024)
by: Robert, Thomas, et al.
Published: (2024)
Multi-Objective Optimization via Wasserstein-Fisher-Rao Gradient Flow
by: Ren, Yinuo, et al.
Published: (2023)
by: Ren, Yinuo, et al.
Published: (2023)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
Greedy Low-Rank Gradient Compression for Distributed Learning with Convergence Guarantees
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
Low-Tubal-Rank Tensor Recovery via Factorized Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2024)
by: Liu, Zhiyu, et al.
Published: (2024)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin
by: Kumar, Akshay, et al.
Published: (2025)
by: Kumar, Akshay, et al.
Published: (2025)
Inclusive KL Minimization: A Wasserstein-Fisher-Rao Gradient Flow Perspective
by: Zhu, Jia-Jie
Published: (2024)
by: Zhu, Jia-Jie
Published: (2024)
Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
by: Oikonomou, Dimitris, et al.
Published: (2025)
by: Oikonomou, Dimitris, et al.
Published: (2025)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
Efficient Low-Tubal-Rank Tensor Estimation via Alternating Preconditioned Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2025)
by: Liu, Zhiyu, et al.
Published: (2025)
Dissipative Gradient Descent Ascent Method: A Control Theory Inspired Algorithm for Min-max Optimization
by: Zheng, Tianqi, et al.
Published: (2024)
by: Zheng, Tianqi, et al.
Published: (2024)
QCQP-Net: Reliably Learning Feasible Alternating Current Optimal Power Flow Solutions Under Constraints
by: Zeng, Sihan, et al.
Published: (2024)
by: Zeng, Sihan, et al.
Published: (2024)
Gradient-Informed Monte Carlo Fine-Tuning of Diffusion Models for Low-Thrust Trajectory Design
by: Graebner, Jannik, et al.
Published: (2025)
by: Graebner, Jannik, et al.
Published: (2025)
Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local Updates
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
by: Kassing, Sebastian, et al.
Published: (2025)
by: Kassing, Sebastian, et al.
Published: (2025)
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel
by: Hagemann, Paul, et al.
Published: (2023)
by: Hagemann, Paul, et al.
Published: (2023)
Neural Wasserstein Gradient Flows for Maximum Mean Discrepancies with Riesz Kernels
by: Altekrüger, Fabian, et al.
Published: (2023)
by: Altekrüger, Fabian, et al.
Published: (2023)
Linear Convergence of Independent Natural Policy Gradient in Games with Entropy Regularization
by: Sun, Youbang, et al.
Published: (2024)
by: Sun, Youbang, et al.
Published: (2024)
Unbiased Gradient Low-Rank Projection
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Reconstructing Physics-Informed Machine Learning for Traffic Flow Modeling: a Multi-Gradient Descent and Pareto Learning Approach
by: Lei, Yuan-Zheng, et al.
Published: (2025)
by: Lei, Yuan-Zheng, et al.
Published: (2025)
Similar Items
-
(How) Learning Rates Regulate Catastrophic Overtraining
by: Rofin, Mark, et al.
Published: (2026) -
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024) -
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024) -
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024) -
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)