Improving Generalization and Convergence by Enhancing Implicit Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Mingze, Wang, Jinbo, He, Haotian, Wang, Zilin, Huang, Guanhua, Xiong, Feiyu, Li, Zhiyu, E, Weinan, Wu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
How Transformers Get Rich: Approximation and Dynamics Analysis
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)
by: Wang, Jinbo, et al.
Published: (2026)
Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
by: Wu, Haotian
Published: (2025)
by: Wu, Haotian
Published: (2025)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Implicit Regularization in Perturbed Deep Matrix Factorization: Spectral Conditions and Stability
by: Wang, Jingzhe, et al.
Published: (2026)
by: Wang, Jingzhe, et al.
Published: (2026)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
by: Cayci, Semih, et al.
Published: (2021)
by: Cayci, Semih, et al.
Published: (2021)
Robust Implicit Regularization via Weight Normalization
by: Chou, Hung-Hsu, et al.
Published: (2023)
by: Chou, Hung-Hsu, et al.
Published: (2023)
Improving Convergence and Generalization Using Parameter Symmetries
by: Zhao, Bo, et al.
Published: (2023)
by: Zhao, Bo, et al.
Published: (2023)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Implicit Bias and Fast Convergence Rates for Self-attention
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence Rates
by: Grazzi, Riccardo, et al.
Published: (2024)
by: Grazzi, Riccardo, et al.
Published: (2024)
Implicit Regularization of Gradient Flow on One-Layer Softmax Attention
by: Sheen, Heejune, et al.
Published: (2024)
by: Sheen, Heejune, et al.
Published: (2024)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Manifold Regularization Classification Model Based On Improved Diffusion Map
by: Guo, Hongfu, et al.
Published: (2024)
by: Guo, Hongfu, et al.
Published: (2024)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
by: Ma, Jianhao, et al.
Published: (2025)
by: Ma, Jianhao, et al.
Published: (2025)
Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness
by: Bo, Yanan, et al.
Published: (2026)
by: Bo, Yanan, et al.
Published: (2026)
Gradient Regularized Newton Boosting Trees with Global Convergence
by: Zozoulenko, Nikita, et al.
Published: (2026)
by: Zozoulenko, Nikita, et al.
Published: (2026)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
Implicit Regularization Makes Overparameterized Asymmetric Matrix Sensing Robust to Perturbations
by: Wind, Johan S.
Published: (2023)
by: Wind, Johan S.
Published: (2023)
Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and Convergence
by: Han, Yinbin, et al.
Published: (2024)
by: Han, Yinbin, et al.
Published: (2024)
Cauchy-Schwarz Regularizers
by: Taner, Sueda, et al.
Published: (2025)
by: Taner, Sueda, et al.
Published: (2025)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
by: Wang, Shengbo
Published: (2026)
by: Wang, Shengbo
Published: (2026)
Regularization for Adversarial Robust Learning
by: Wang, Jie, et al.
Published: (2024)
by: Wang, Jie, et al.
Published: (2024)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
On Generalization and Regularization via Wasserstein Distributionally Robust Optimization
by: Wu, Qinyu, et al.
Published: (2022)
by: Wu, Qinyu, et al.
Published: (2022)
Distributed Online Convex Optimization with Nonseparable Costs and Constraints
by: Pan, Zhaoye, et al.
Published: (2026)
by: Pan, Zhaoye, et al.
Published: (2026)
Communication Efficient Federated Learning with Linear Convergence on Heterogeneous Data
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Divergence Results and Convergence of a Variance Reduced Version of ADAM
by: Wang, Ruiqi, et al.
Published: (2022)
by: Wang, Ruiqi, et al.
Published: (2022)
Provably Convergent Federated Trilevel Learning
by: Jiao, Yang, et al.
Published: (2023)
by: Jiao, Yang, et al.
Published: (2023)
Enhancing Distributional Robustness in Principal Component Analysis by Wasserstein Distances
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Inertial Quadratic Majorization Minimization with Application to Kernel Regularized Learning
by: Heng, Qiang, et al.
Published: (2025)
by: Heng, Qiang, et al.
Published: (2025)
Distributed Random Reshuffling Methods with Improved Convergence
by: Huang, Kun, et al.
Published: (2023)
by: Huang, Kun, et al.
Published: (2023)
Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Similar Items
-
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025) -
How Transformers Get Rich: Approximation and Dynamics Analysis
by: Wang, Mingze, et al.
Published: (2024) -
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025) -
Achieving Margin Maximization Exponentially Fast via Progressive Norm Rescaling
by: Wang, Mingze, et al.
Published: (2023) -
Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws
by: Wang, Jinbo, et al.
Published: (2026)