Correction of Decoupled Weight Decay
Fuente:
arXiv
Salvato in:
| Autore principale: | Chou, Jason Chuan-Chih |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
di: Chooi, Jay, et al.
Pubblicazione: (2025)
di: Chooi, Jay, et al.
Pubblicazione: (2025)
Embedding Geometries of Contrastive Language-Image Pre-Training
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2024)
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2024)
ViT Registers and Fractal ViT
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2026)
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2026)
Decoupled Weight Decay for Any $p$ Norm
di: Outmezguine, Nadav Joseph, et al.
Pubblicazione: (2024)
di: Outmezguine, Nadav Joseph, et al.
Pubblicazione: (2024)
Cautious Weight Decay
di: Chen, Lizhang, et al.
Pubblicazione: (2025)
di: Chen, Lizhang, et al.
Pubblicazione: (2025)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
di: Fu, Xiaoliang, et al.
Pubblicazione: (2026)
di: Fu, Xiaoliang, et al.
Pubblicazione: (2026)
Does Weight Decay Enhance Training Stability?
di: Saether, Marius, et al.
Pubblicazione: (2026)
di: Saether, Marius, et al.
Pubblicazione: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
Towards Understanding Neural Collapse: The Effects of Batch Normalization and Weight Decay
di: Pan, Leyan, et al.
Pubblicazione: (2023)
di: Pan, Leyan, et al.
Pubblicazione: (2023)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
di: Galanti, Tomer, et al.
Pubblicazione: (2022)
di: Galanti, Tomer, et al.
Pubblicazione: (2022)
The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression
di: Xu, Yongzhong
Pubblicazione: (2026)
di: Xu, Yongzhong
Pubblicazione: (2026)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
di: He, Di, et al.
Pubblicazione: (2025)
di: He, Di, et al.
Pubblicazione: (2025)
Cosine-Gated Adam-Decay: Drop-In Staleness-Aware Outer Optimization for Decoupled DiLoCo
di: Shah, Vatsal, et al.
Pubblicazione: (2026)
di: Shah, Vatsal, et al.
Pubblicazione: (2026)
Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
di: Zhang, Kaiqi, et al.
Pubblicazione: (2022)
di: Zhang, Kaiqi, et al.
Pubblicazione: (2022)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
di: Kosson, Atli, et al.
Pubblicazione: (2023)
di: Kosson, Atli, et al.
Pubblicazione: (2023)
Learning to Forget: Continual Learning with Adaptive Weight Decay
di: Ramesh, Aditya A., et al.
Pubblicazione: (2026)
di: Ramesh, Aditya A., et al.
Pubblicazione: (2026)
On the Role of Weight Decay in Collaborative Filtering: A Popularity Perspective
di: Loveland, Donald, et al.
Pubblicazione: (2025)
di: Loveland, Donald, et al.
Pubblicazione: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Fan, Zhiyuan, et al.
Pubblicazione: (2025)
Weighted Temporal Decay Loss for Learning Wearable PPG Data with Sparse Clinical Labels
di: Chung, Yunsung, et al.
Pubblicazione: (2026)
di: Chung, Yunsung, et al.
Pubblicazione: (2026)
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
di: Chen, Ke, et al.
Pubblicazione: (2024)
di: Chen, Ke, et al.
Pubblicazione: (2024)
OUIDecay: Adaptive Layer-wise Weight Decay for CNNs Using Online Activation Patterns
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
A Self-Attentive Meta-Optimizer with Group-Adaptive Learning Rates and Weight Decay
di: Zhao, JiangBo, et al.
Pubblicazione: (2026)
di: Zhao, JiangBo, et al.
Pubblicazione: (2026)
Robust Implicit Regularization via Weight Normalization
di: Chou, Hung-Hsu, et al.
Pubblicazione: (2023)
di: Chou, Hung-Hsu, et al.
Pubblicazione: (2023)
Constraint Learning in Multi-Agent Dynamic Games from Demonstrations of Local Nash Interactions
di: Zhang, Zhouyu, et al.
Pubblicazione: (2025)
di: Zhang, Zhouyu, et al.
Pubblicazione: (2025)
Decoupled Travel Planning with Behavior Forest
di: Yuan, Duanyang, et al.
Pubblicazione: (2026)
di: Yuan, Duanyang, et al.
Pubblicazione: (2026)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
di: Kosson, Atli, et al.
Pubblicazione: (2025)
di: Kosson, Atli, et al.
Pubblicazione: (2025)
SG-XDEAT: Sparsity-Guided Cross-Dimensional and Cross-Encoding Attention with Target-Aware Conditioning in Tabular Learning
di: Cheng, Chih-Chuan, et al.
Pubblicazione: (2025)
di: Cheng, Chih-Chuan, et al.
Pubblicazione: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
di: Jacot, Arthur, et al.
Pubblicazione: (2024)
di: Jacot, Arthur, et al.
Pubblicazione: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
di: Xu, Huangyu, et al.
Pubblicazione: (2026)
di: Xu, Huangyu, et al.
Pubblicazione: (2026)
The Rank and Gradient Lost in Non-stationarity: Sample Weight Decay for Mitigating Plasticity Loss in Reinforcement Learning
di: Wu, Zihao, et al.
Pubblicazione: (2026)
di: Wu, Zihao, et al.
Pubblicazione: (2026)
Online Imitation Learning for Manipulation via Decaying Relative Correction through Teleoperation
di: Pan, Cheng, et al.
Pubblicazione: (2025)
di: Pan, Cheng, et al.
Pubblicazione: (2025)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
di: Xu, Yongzhong
Pubblicazione: (2026)
di: Xu, Yongzhong
Pubblicazione: (2026)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
di: Xie, Zeke, et al.
Pubblicazione: (2020)
di: Xie, Zeke, et al.
Pubblicazione: (2020)
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
di: Verma, Lucky
Pubblicazione: (2026)
di: Verma, Lucky
Pubblicazione: (2026)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
di: Tian, Junjiao, et al.
Pubblicazione: (2024)
di: Tian, Junjiao, et al.
Pubblicazione: (2024)
Tune without Validation: Searching for Learning Rate and Weight Decay on Training Sets
di: Brigato, Lorenzo, et al.
Pubblicazione: (2024)
di: Brigato, Lorenzo, et al.
Pubblicazione: (2024)
Trainable Weight Averaging: Accelerating Training and Improving Generalization
di: Li, Tao, et al.
Pubblicazione: (2022)
di: Li, Tao, et al.
Pubblicazione: (2022)
Impact of Loss Weight and Model Complexity on Physics-Informed Neural Networks for Computational Fluid Dynamics
di: Chou, Yi En, et al.
Pubblicazione: (2025)
di: Chou, Yi En, et al.
Pubblicazione: (2025)
AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training
di: Guo, Fu-Ming, et al.
Pubblicazione: (2025)
di: Guo, Fu-Ming, et al.
Pubblicazione: (2025)
Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization
di: Das, Abhijit, et al.
Pubblicazione: (2026)
di: Das, Abhijit, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
di: Chooi, Jay, et al.
Pubblicazione: (2025) -
Embedding Geometries of Contrastive Language-Image Pre-Training
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2024) -
ViT Registers and Fractal ViT
di: Chou, Jason Chuan-Chih, et al.
Pubblicazione: (2026) -
Decoupled Weight Decay for Any $p$ Norm
di: Outmezguine, Nadav Joseph, et al.
Pubblicazione: (2024) -
Cautious Weight Decay
di: Chen, Lizhang, et al.
Pubblicazione: (2025)