On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xie, Zeke, Xu, Zhiqiang, Zhang, Jingzhao, Sato, Issei, Sugiyama, Masashi |
|---|---|
| Format: | Preprint |
| Publié: |
2020
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Understanding Transformer Optimization via Gradient Heterogeneity
par: Tomihari, Akiyoshi, et autres
Publié: (2025)
par: Tomihari, Akiyoshi, et autres
Publié: (2025)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
par: Xu, Kevin, et autres
Publié: (2025)
par: Xu, Kevin, et autres
Publié: (2025)
A Formal Comparison Between Chain of Thought and Latent Thought
par: Xu, Kevin, et autres
Publié: (2025)
par: Xu, Kevin, et autres
Publié: (2025)
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
par: Chen, Lesi, et autres
Publié: (2023)
par: Chen, Lesi, et autres
Publié: (2023)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
par: Tanaka, Yuto, et autres
Publié: (2026)
par: Tanaka, Yuto, et autres
Publié: (2026)
Decoupled Weight Decay for Any $p$ Norm
par: Outmezguine, Nadav Joseph, et autres
Publié: (2024)
par: Outmezguine, Nadav Joseph, et autres
Publié: (2024)
Learning Robust Diffusion Models from Imprecise Supervision
par: Wu, Dong-Dong, et autres
Publié: (2025)
par: Wu, Dong-Dong, et autres
Publié: (2025)
VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction
par: Cai, Xin-Qiang, et autres
Publié: (2026)
par: Cai, Xin-Qiang, et autres
Publié: (2026)
GradientStabilizer:Fix the Norm, Not the Gradient
par: Huang, Tianjin, et autres
Publié: (2025)
par: Huang, Tianjin, et autres
Publié: (2025)
Random Masking Finds Winning Tickets for Parameter Efficient Fine-tuning
par: Xu, Jing, et autres
Publié: (2024)
par: Xu, Jing, et autres
Publié: (2024)
Mano: Restriking Manifold Optimization for LLM Training
par: Gu, Yufei, et autres
Publié: (2026)
par: Gu, Yufei, et autres
Publié: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
par: Ackermann, Johannes, et autres
Publié: (2026)
par: Ackermann, Johannes, et autres
Publié: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
par: Ono, Shinnosuke, et autres
Publié: (2026)
par: Ono, Shinnosuke, et autres
Publié: (2026)
Offline Reinforcement Learning from Datasets with Structured Non-Stationarity
par: Ackermann, Johannes, et autres
Publié: (2024)
par: Ackermann, Johannes, et autres
Publié: (2024)
Low Rank Gradients and Where to Find Them
par: Sonthalia, Rishi, et autres
Publié: (2025)
par: Sonthalia, Rishi, et autres
Publié: (2025)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
par: Nishimori, Soichiro, et autres
Publié: (2025)
par: Nishimori, Soichiro, et autres
Publié: (2025)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
par: Gao, Chengqian, et autres
Publié: (2025)
par: Gao, Chengqian, et autres
Publié: (2025)
On the Condition Number Dependency in Bilevel Optimization
par: Chen, Lesi, et autres
Publié: (2025)
par: Chen, Lesi, et autres
Publié: (2025)
Reasoning Inconsistencies and How to Mitigate Them in Deep Learning
par: Arakelyan, Erik
Publié: (2025)
par: Arakelyan, Erik
Publié: (2025)
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
par: Hassanpour, Negar, et autres
Publié: (2025)
par: Hassanpour, Negar, et autres
Publié: (2025)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
par: Xu, Huangyu, et autres
Publié: (2026)
par: Xu, Huangyu, et autres
Publié: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
par: Ackermann, Johannes, et autres
Publié: (2025)
par: Ackermann, Johannes, et autres
Publié: (2025)
Towards Scalable Oversight via Partitioned Human Supervision
par: Yin, Ren, et autres
Publié: (2025)
par: Yin, Ren, et autres
Publié: (2025)
Weak-to-Strong Diffusion with Reflection
par: Bai, Lichen, et autres
Publié: (2025)
par: Bai, Lichen, et autres
Publié: (2025)
The Geometry of Multi-Task Grokking: Transverse Instability, Superposition, and Weight Decay Phase Structure
par: Xu, Yongzhong
Publié: (2026)
par: Xu, Yongzhong
Publié: (2026)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
par: Fu, Xiaoliang, et autres
Publié: (2026)
par: Fu, Xiaoliang, et autres
Publié: (2026)
GNN Explanations that do not Explain and How to find Them
par: Azzolin, Steve, et autres
Publié: (2026)
par: Azzolin, Steve, et autres
Publié: (2026)
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
par: Chen, Hao, et autres
Publié: (2023)
par: Chen, Hao, et autres
Publié: (2023)
How to Square Tensor Networks and Circuits Without Squaring Them
par: Loconte, Lorenzo, et autres
Publié: (2025)
par: Loconte, Lorenzo, et autres
Publié: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
par: Huang, Jerry, et autres
Publié: (2025)
par: Huang, Jerry, et autres
Publié: (2025)
Sharpness-Aware Black-Box Optimization
par: Ye, Feiyang, et autres
Publié: (2024)
par: Ye, Feiyang, et autres
Publié: (2024)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
par: Xu, Haoran, et autres
Publié: (2025)
par: Xu, Haoran, et autres
Publié: (2025)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
par: Fan, Zhiyuan, et autres
Publié: (2025)
par: Fan, Zhiyuan, et autres
Publié: (2025)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
par: Wen, Xuexiang, et autres
Publié: (2026)
par: Wen, Xuexiang, et autres
Publié: (2026)
One Wave To Explain Them All: A Unifying Perspective On Feature Attribution
par: Kasmi, Gabriel, et autres
Publié: (2024)
par: Kasmi, Gabriel, et autres
Publié: (2024)
Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
par: Prinster, Drew, et autres
Publié: (2024)
par: Prinster, Drew, et autres
Publié: (2024)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
par: He, Di, et autres
Publié: (2025)
par: He, Di, et autres
Publié: (2025)
Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
par: Nishimori, Soichiro, et autres
Publié: (2026)
par: Nishimori, Soichiro, et autres
Publié: (2026)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
par: Liu, Siyuan, et autres
Publié: (2026)
par: Liu, Siyuan, et autres
Publié: (2026)
The Pitfalls of KV Cache Compression
par: Chen, Alex, et autres
Publié: (2025)
par: Chen, Alex, et autres
Publié: (2025)
Documents similaires
-
Understanding Transformer Optimization via Gradient Heterogeneity
par: Tomihari, Akiyoshi, et autres
Publié: (2025) -
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
par: Xu, Kevin, et autres
Publié: (2025) -
A Formal Comparison Between Chain of Thought and Latent Thought
par: Xu, Kevin, et autres
Publié: (2025) -
On Finding Small Hyper-Gradients in Bilevel Optimization: Hardness Results and Improved Analysis
par: Chen, Lesi, et autres
Publié: (2023) -
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
par: Tanaka, Yuto, et autres
Publié: (2026)