Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
Fuente:
arXiv
Saved in:
| Main Author: | Abouzeid, Abdelrahman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
BNPO: Beta Normalization Policy Optimization
by: Xiao, Changyi, et al.
Published: (2025)
by: Xiao, Changyi, et al.
Published: (2025)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
by: Scetbon, Meyer, et al.
Published: (2025)
by: Scetbon, Meyer, et al.
Published: (2025)
Continuous-Time Analysis of Adaptive Optimization and Normalization
by: Gould, Rhys, et al.
Published: (2024)
by: Gould, Rhys, et al.
Published: (2024)
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Latent Bayesian Optimization via Autoregressive Normalizing Flows
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
Optimization and Generalization Guarantees for Weight Normalization
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
How to Train Your LLM Web Agent: A Statistical Diagnosis
by: Vattikonda, Dheeraj, et al.
Published: (2025)
by: Vattikonda, Dheeraj, et al.
Published: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Robust Optimization in Causal Models and G-Causal Normalizing Flows
by: Visentin, Gabriele, et al.
Published: (2025)
by: Visentin, Gabriele, et al.
Published: (2025)
AlphaGrad: Non-Linear Gradient Normalization Optimizer
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
Efficient Regression-Based Training of Normalizing Flows for Boltzmann Generators
by: Rehman, Danyal, et al.
Published: (2025)
by: Rehman, Danyal, et al.
Published: (2025)
Mitigating Gradient Overlap in Deep Residual Networks with Gradient Normalization for Improved Non-Convex Optimization
by: Yun, Juyoung
Published: (2024)
by: Yun, Juyoung
Published: (2024)
Conda: Column-Normalized Adam for Training Large Language Models Faster
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Mano: Restriking Manifold Optimization for LLM Training
by: Gu, Yufei, et al.
Published: (2026)
by: Gu, Yufei, et al.
Published: (2026)
On the Nonlinearity of Layer Normalization
by: Ni, Yunhao, et al.
Published: (2024)
by: Ni, Yunhao, et al.
Published: (2024)
NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
by: Tarasov, Denis, et al.
Published: (2025)
by: Tarasov, Denis, et al.
Published: (2025)
Group-in-Group Policy Optimization for LLM Agent Training
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
by: Zhang, Huishuai, et al.
Published: (2025)
by: Zhang, Huishuai, et al.
Published: (2025)
Non-Normal Diffusion Models
by: Li, Henry
Published: (2024)
by: Li, Henry
Published: (2024)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
by: Fluri, Lukas, et al.
Published: (2024)
by: Fluri, Lukas, et al.
Published: (2024)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Learning Rate Transfer in Normalized Transformers
by: Shigida, Boris, et al.
Published: (2026)
by: Shigida, Boris, et al.
Published: (2026)
Scaling CrossQ with Weight Normalization
by: Palenicek, Daniel, et al.
Published: (2025)
by: Palenicek, Daniel, et al.
Published: (2025)
On the Weight Dynamics of Deep Normalized Networks
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2023)
by: Mehmeti-Göpel, Christian H. X. Ali, et al.
Published: (2023)
Amortized Sampling with Transferable Normalizing Flows
by: Tan, Charlie B., et al.
Published: (2025)
by: Tan, Charlie B., et al.
Published: (2025)
ANAct: Adaptive Normalization for Activation Functions
by: Peiwen, Yuan, et al.
Published: (2022)
by: Peiwen, Yuan, et al.
Published: (2022)
Knowledge Graph Embedding by Normalizing Flows
by: Xiao, Changyi, et al.
Published: (2024)
by: Xiao, Changyi, et al.
Published: (2024)
Batch Normalization Amplifies Memorization and Privacy Risks
by: Doan, Ngoc Phu, et al.
Published: (2026)
by: Doan, Ngoc Phu, et al.
Published: (2026)
Inner-Instance Normalization for Time Series Forecasting
by: Jibao, Zipo, et al.
Published: (2025)
by: Jibao, Zipo, et al.
Published: (2025)
Riemannian Batch Normalization: A Gyro Approach
by: Chen, Ziheng, et al.
Published: (2025)
by: Chen, Ziheng, et al.
Published: (2025)
GRANOLA: Adaptive Normalization for Graph Neural Networks
by: Eliasof, Moshe, et al.
Published: (2024)
by: Eliasof, Moshe, et al.
Published: (2024)
Disentangling Neural Disjunctive Normal Form Models
by: Baugh, Kexin Gu, et al.
Published: (2025)
by: Baugh, Kexin Gu, et al.
Published: (2025)
Channel Normalization for Time Series Channel Identification
by: Lee, Seunghan, et al.
Published: (2025)
by: Lee, Seunghan, et al.
Published: (2025)
Self-Normalized Resets for Plasticity in Continual Learning
by: Farias, Vivek F., et al.
Published: (2024)
by: Farias, Vivek F., et al.
Published: (2024)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
by: Cheng, Pengyu, et al.
Published: (2023)
by: Cheng, Pengyu, et al.
Published: (2023)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
A Comparative Study on How Data Normalization Affects Zero-Shot Generalization in Time Series Foundation Models
by: Ahmed, Ihab, et al.
Published: (2025)
by: Ahmed, Ihab, et al.
Published: (2025)
Similar Items
-
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026) -
BNPO: Beta Normalization Policy Optimization
by: Xiao, Changyi, et al.
Published: (2025) -
Gradient Multi-Normalization for Stateless and Scalable LLM Training
by: Scetbon, Meyer, et al.
Published: (2025) -
Continuous-Time Analysis of Adaptive Optimization and Normalization
by: Gould, Rhys, et al.
Published: (2024) -
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
by: Ma, Chao, et al.
Published: (2024)