Optimization and Generalization Guarantees for Weight Normalization
Fuente:
arXiv
Saved in:
| Main Authors: | Cisneros-Velarde, Pedro, Chen, Zhijie, Koyejo, Sanmi, Banerjee, Arindam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimization for Neural Operators can Benefit from Width
by: Cisneros-Velarde, Pedro, et al.
Published: (2025)
by: Cisneros-Velarde, Pedro, et al.
Published: (2025)
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
by: Ma, Jianhao, et al.
Published: (2025)
by: Ma, Jianhao, et al.
Published: (2025)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
by: Schotthöfer, Steffen, et al.
Published: (2024)
by: Schotthöfer, Steffen, et al.
Published: (2024)
Optimal Control Operator Perspective and a Neural Adaptive Spectral Method
by: Feng, Mingquan, et al.
Published: (2024)
by: Feng, Mingquan, et al.
Published: (2024)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
by: Li, Gang, et al.
Published: (2024)
by: Li, Gang, et al.
Published: (2024)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
by: Farhi, Nadir
Published: (2025)
by: Farhi, Nadir
Published: (2025)
On the Condition Number Dependency in Bilevel Optimization
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
by: C., Simo Alami, et al.
Published: (2025)
by: C., Simo Alami, et al.
Published: (2025)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
Multi-CALF: A Policy Combination Approach with Statistical Guarantees
by: Malaniya, Georgiy, et al.
Published: (2025)
by: Malaniya, Georgiy, et al.
Published: (2025)
Personalized Multi-tier Federated Learning
by: Banerjee, Sourasekhar, et al.
Published: (2024)
by: Banerjee, Sourasekhar, et al.
Published: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
Time-Varying Optimization for Streaming Data Via Temporal Weighting
by: Abrar, Muhammad Faraz Ul, et al.
Published: (2025)
by: Abrar, Muhammad Faraz Ul, et al.
Published: (2025)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
by: Allouah, Youssef, et al.
Published: (2024)
by: Allouah, Youssef, et al.
Published: (2024)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
Scalable Data-Driven Reachability Analysis and Control via Koopman Operators with Conformal Coverage Guarantees
by: Nath, Devesh, et al.
Published: (2026)
by: Nath, Devesh, et al.
Published: (2026)
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime
by: Geng, Haoyu, et al.
Published: (2023)
by: Geng, Haoyu, et al.
Published: (2023)
A Safe Screening Rule with Bi-level Optimization of $ν$ Support Vector Machine
by: Yang, Zhiji, et al.
Published: (2024)
by: Yang, Zhiji, et al.
Published: (2024)
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
by: Sharifnassab, Arsalan, et al.
Published: (2024)
by: Sharifnassab, Arsalan, et al.
Published: (2024)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
by: Wasserkrug, Segev, et al.
Published: (2024)
by: Wasserkrug, Segev, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Jacobian Descent for Multi-Objective Optimization
by: Quinton, Pierre, et al.
Published: (2024)
by: Quinton, Pierre, et al.
Published: (2024)
Neural Solver Selection for Combinatorial Optimization
by: Gao, Chengrui, et al.
Published: (2024)
by: Gao, Chengrui, et al.
Published: (2024)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
Budget-aware Auto Optimizer Configurator
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Client-Centric Federated Adaptive Optimization
by: Sun, Jianhui, et al.
Published: (2025)
by: Sun, Jianhui, et al.
Published: (2025)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
by: Onorato, Gabriele
Published: (2024)
by: Onorato, Gabriele
Published: (2024)
Similar Items
-
Optimization for Neural Operators can Benefit from Width
by: Cisneros-Velarde, Pedro, et al.
Published: (2025) -
Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
by: Ma, Jianhao, et al.
Published: (2025) -
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2023) -
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024) -
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
by: Schotthöfer, Steffen, et al.
Published: (2024)