Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Kosson, Atli, Messmer, Bettina, Jaggi, Martin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
di: Kosson, Atli, et al.
Pubblicazione: (2024)
di: Kosson, Atli, et al.
Pubblicazione: (2024)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
di: Kosson, Atli, et al.
Pubblicazione: (2025)
di: Kosson, Atli, et al.
Pubblicazione: (2025)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
di: Messmer, Bettina, et al.
Pubblicazione: (2025)
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
di: Fan, Dongyang, et al.
Pubblicazione: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
di: Fan, Simin, et al.
Pubblicazione: (2024)
di: Fan, Simin, et al.
Pubblicazione: (2024)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
di: Galanti, Tomer, et al.
Pubblicazione: (2022)
di: Galanti, Tomer, et al.
Pubblicazione: (2022)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
di: Jacot, Arthur, et al.
Pubblicazione: (2024)
di: Jacot, Arthur, et al.
Pubblicazione: (2024)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
di: He, Di, et al.
Pubblicazione: (2025)
di: He, Di, et al.
Pubblicazione: (2025)
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
di: Chen, Ke, et al.
Pubblicazione: (2024)
di: Chen, Ke, et al.
Pubblicazione: (2024)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
di: Pagliardini, Matteo, et al.
Pubblicazione: (2024)
di: Pagliardini, Matteo, et al.
Pubblicazione: (2024)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
di: Chayti, El Mahdi, et al.
Pubblicazione: (2024)
di: Chayti, El Mahdi, et al.
Pubblicazione: (2024)
Fair Bilevel Neural Network (FairBiNN): On Balancing fairness and accuracy via Stackelberg Equilibrium
di: Yazdani-Jahromi, Mehdi, et al.
Pubblicazione: (2024)
di: Yazdani-Jahromi, Mehdi, et al.
Pubblicazione: (2024)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
di: Pang, Tianyu, et al.
Pubblicazione: (2026)
Towards Understanding Neural Collapse: The Effects of Batch Normalization and Weight Decay
di: Pan, Leyan, et al.
Pubblicazione: (2023)
di: Pan, Leyan, et al.
Pubblicazione: (2023)
Dendritic Neural Networks with Equilibrium Propagation
di: Kubo, Yoshimasa
Pubblicazione: (2026)
di: Kubo, Yoshimasa
Pubblicazione: (2026)
Correction of Decoupled Weight Decay
di: Chou, Jason Chuan-Chih
Pubblicazione: (2025)
di: Chou, Jason Chuan-Chih
Pubblicazione: (2025)
Cautious Weight Decay
di: Chen, Lizhang, et al.
Pubblicazione: (2025)
di: Chen, Lizhang, et al.
Pubblicazione: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
di: Turki, Yassine, et al.
Pubblicazione: (2026)
di: Turki, Yassine, et al.
Pubblicazione: (2026)
Learning to Forget: Continual Learning with Adaptive Weight Decay
di: Ramesh, Aditya A., et al.
Pubblicazione: (2026)
di: Ramesh, Aditya A., et al.
Pubblicazione: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression
di: Xu, Yongzhong
Pubblicazione: (2026)
di: Xu, Yongzhong
Pubblicazione: (2026)
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
di: Zhou, Xinyu, et al.
Pubblicazione: (2024)
di: Zhou, Xinyu, et al.
Pubblicazione: (2024)
Benchmarking Optimizers for Large Language Model Pretraining
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
CoBo: Collaborative Learning via Bilevel Optimization
di: Hashemi, Diba, et al.
Pubblicazione: (2024)
di: Hashemi, Diba, et al.
Pubblicazione: (2024)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
di: Laus, Hannah, et al.
Pubblicazione: (2025)
di: Laus, Hannah, et al.
Pubblicazione: (2025)
Using Machine Learning for move sequence visualization and generation in climbing
di: Rimbot, Thomas, et al.
Pubblicazione: (2025)
di: Rimbot, Thomas, et al.
Pubblicazione: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
Stochastic Difference-of-Convex Optimization with Momentum
di: Chayti, El Mahdi, et al.
Pubblicazione: (2025)
di: Chayti, El Mahdi, et al.
Pubblicazione: (2025)
A Split-Client Approach to Second-Order Optimization
di: Chayti, El Mahdi, et al.
Pubblicazione: (2025)
di: Chayti, El Mahdi, et al.
Pubblicazione: (2025)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
di: Xie, Zeke, et al.
Pubblicazione: (2020)
di: Xie, Zeke, et al.
Pubblicazione: (2020)
Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
di: Zhang, Kaiqi, et al.
Pubblicazione: (2022)
di: Zhang, Kaiqi, et al.
Pubblicazione: (2022)
Randomly Weighted Neuromodulation in Neural Networks Facilitates Learning of Manifolds Common Across Tasks
di: Hong, Jinyung, et al.
Pubblicazione: (2023)
di: Hong, Jinyung, et al.
Pubblicazione: (2023)
How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks
di: Frati, Lapo, et al.
Pubblicazione: (2025)
di: Frati, Lapo, et al.
Pubblicazione: (2025)
Bridging Equilibrium and Kinetics Prediction with a Data-Weighted Neural Network Model of Methane Steam Reforming
di: Pizoń, Zofia, et al.
Pubblicazione: (2025)
di: Pizoń, Zofia, et al.
Pubblicazione: (2025)
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
di: Fan, Simin, et al.
Pubblicazione: (2025)
di: Fan, Simin, et al.
Pubblicazione: (2025)
Weighted Temporal Decay Loss for Learning Wearable PPG Data with Sparse Clinical Labels
di: Chung, Yunsung, et al.
Pubblicazione: (2026)
di: Chung, Yunsung, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
di: Kosson, Atli, et al.
Pubblicazione: (2024) -
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
di: Kosson, Atli, et al.
Pubblicazione: (2025) -
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
di: Dremov, Aleksandr, et al.
Pubblicazione: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
di: Messmer, Bettina, et al.
Pubblicazione: (2025) -
Towards an empirical understanding of MoE design choices
di: Fan, Dongyang, et al.
Pubblicazione: (2024)