SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Chao, Gong, Wenbo, Scetbon, Meyer, Meeds, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gradient Multi-Normalization for Stateless and Scalable LLM Training
by: Scetbon, Meyer, et al.
Published: (2025)
by: Scetbon, Meyer, et al.
Published: (2025)
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
by: Gong, Wenbo, et al.
Published: (2025)
by: Gong, Wenbo, et al.
Published: (2025)
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025)
by: Saraswatula, Ashwin, et al.
Published: (2025)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)
by: Gupta, Tarun, et al.
Published: (2024)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
by: Abouzeid, Abdelrahman
Published: (2026)
by: Abouzeid, Abdelrahman
Published: (2026)
DC-SGD: Differentially Private SGD with Dynamic Clipping through Gradient Norm Distribution Estimation
by: Wei, Chengkun, et al.
Published: (2025)
by: Wei, Chengkun, et al.
Published: (2025)
Comprehensive Forecasting-Based Analysis of Hybrid and Stacked Stateful/ Stateless Models
by: Saha, Swayamjit
Published: (2024)
by: Saha, Swayamjit
Published: (2024)
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
by: Sun, Ruotong, et al.
Published: (2026)
by: Sun, Ruotong, et al.
Published: (2026)
Whitening Not Recommended for Classification Tasks in LLMs
by: Forooghi, Ali, et al.
Published: (2024)
by: Forooghi, Ali, et al.
Published: (2024)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
by: Yu, Dingzhi, et al.
Published: (2026)
by: Yu, Dingzhi, et al.
Published: (2026)
Reservoir Subspace Injection for Online ICA under Top-n Whitening
by: Xiao, Wenjun, et al.
Published: (2026)
by: Xiao, Wenjun, et al.
Published: (2026)
LeanTTA: A Backpropagation-Free and Stateless Approach to Quantized Test-Time Adaptation on Edge Devices
by: Dong, Cynthia, et al.
Published: (2025)
by: Dong, Cynthia, et al.
Published: (2025)
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024)
by: Christmann, Andreas, et al.
Published: (2024)
Low-Rank Correction for Quantized LLMs
by: Scetbon, Meyer, et al.
Published: (2024)
by: Scetbon, Meyer, et al.
Published: (2024)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
Accumulative SGD Influence Estimation for Data Attribution
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Neural Structure Learning with Stochastic Differential Equations
by: Wang, Benjie, et al.
Published: (2023)
by: Wang, Benjie, et al.
Published: (2023)
Counter-Samples: A Stateless Strategy to Neutralize Black Box Adversarial Attacks
by: Bokobza, Roey, et al.
Published: (2024)
by: Bokobza, Roey, et al.
Published: (2024)
ARO: A New Lens On Matrix Optimization For Large Models
by: Gong, Wenbo, et al.
Published: (2026)
by: Gong, Wenbo, et al.
Published: (2026)
Dr. Post-Training: A Data Regularization Perspective on LLM Post-Training
by: Hu, Pingbang, et al.
Published: (2026)
by: Hu, Pingbang, et al.
Published: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Worker Disagreement Reveals Sharp Directions in Local SGD
by: Dimlioglu, Tolga, et al.
Published: (2026)
by: Dimlioglu, Tolga, et al.
Published: (2026)
Transformers are Stateless Differentiable Neural Computers
by: Tang, Bo, et al.
Published: (2026)
by: Tang, Bo, et al.
Published: (2026)
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
RQP-SGD: Differential Private Machine Learning through Noisy SGD and Randomized Quantization
by: Feng, Ce, et al.
Published: (2024)
by: Feng, Ce, et al.
Published: (2024)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization
by: Lei, Yunwen, et al.
Published: (2023)
by: Lei, Yunwen, et al.
Published: (2023)
Efficient Regression-Based Training of Normalizing Flows for Boltzmann Generators
by: Rehman, Danyal, et al.
Published: (2025)
by: Rehman, Danyal, et al.
Published: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
E^2-LLM: Bridging Neural Signals and Interpretable Affective Analysis
by: Ma, Fei, et al.
Published: (2026)
by: Ma, Fei, et al.
Published: (2026)
Mixed-Sample SGD: an End-to-end Analysis of Supervised Transfer Learning
by: Deng, Yuyang, et al.
Published: (2025)
by: Deng, Yuyang, et al.
Published: (2025)
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
Making Batch Normalization Great in Federated Deep Learning
by: Zhong, Jike, et al.
Published: (2023)
by: Zhong, Jike, et al.
Published: (2023)
Conda: Column-Normalized Adam for Training Large Language Models Faster
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training
by: Gong, Xue, et al.
Published: (2026)
by: Gong, Xue, et al.
Published: (2026)
Enhancing DP-SGD through Non-monotonous Adaptive Scaling Gradient Weight
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Similar Items
-
Gradient Multi-Normalization for Stateless and Scalable LLM Training
by: Scetbon, Meyer, et al.
Published: (2025) -
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
by: Gong, Wenbo, et al.
Published: (2025) -
Data Whitening Improves Sparse Autoencoder Learning
by: Saraswatula, Ashwin, et al.
Published: (2025) -
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026) -
The Essential Role of Causality in Foundation World Models for Embodied AI
by: Gupta, Tarun, et al.
Published: (2024)