Gespeichert in:
| Hauptverfasser: | Kosson, Atli, Messmer, Bettina, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2305.17212 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
von: Kosson, Atli, et al.
Veröffentlicht: (2024)
von: Kosson, Atli, et al.
Veröffentlicht: (2024)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
von: Kosson, Atli, et al.
Veröffentlicht: (2025)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
von: Hägele, Alexander, et al.
Veröffentlicht: (2024)
von: Hägele, Alexander, et al.
Veröffentlicht: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
von: Fan, Simin, et al.
Veröffentlicht: (2024)
von: Fan, Simin, et al.
Veröffentlicht: (2024)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
von: He, Di, et al.
Veröffentlicht: (2025)
von: He, Di, et al.
Veröffentlicht: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
von: Pagliardini, Matteo, et al.
Veröffentlicht: (2024)
von: Pagliardini, Matteo, et al.
Veröffentlicht: (2024)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
von: Turki, Yassine, et al.
Veröffentlicht: (2026)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2024)
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2024)
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
von: Chen, Ke, et al.
Veröffentlicht: (2024)
von: Chen, Ke, et al.
Veröffentlicht: (2024)
Fair Bilevel Neural Network (FairBiNN): On Balancing fairness and accuracy via Stackelberg Equilibrium
von: Yazdani-Jahromi, Mehdi, et al.
Veröffentlicht: (2024)
von: Yazdani-Jahromi, Mehdi, et al.
Veröffentlicht: (2024)
Towards Understanding Neural Collapse: The Effects of Batch Normalization and Weight Decay
von: Pan, Leyan, et al.
Veröffentlicht: (2023)
von: Pan, Leyan, et al.
Veröffentlicht: (2023)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
von: Pang, Tianyu, et al.
Veröffentlicht: (2026)
CoBo: Collaborative Learning via Bilevel Optimization
von: Hashemi, Diba, et al.
Veröffentlicht: (2024)
von: Hashemi, Diba, et al.
Veröffentlicht: (2024)
Correction of Decoupled Weight Decay
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
von: Chou, Jason Chuan-Chih
Veröffentlicht: (2025)
Dendritic Neural Networks with Equilibrium Propagation
von: Kubo, Yoshimasa
Veröffentlicht: (2026)
von: Kubo, Yoshimasa
Veröffentlicht: (2026)
Using Machine Learning for move sequence visualization and generation in climbing
von: Rimbot, Thomas, et al.
Veröffentlicht: (2025)
von: Rimbot, Thomas, et al.
Veröffentlicht: (2025)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2024)
Benchmarking Optimizers for Large Language Model Pretraining
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
Stochastic Difference-of-Convex Optimization with Momentum
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2025)
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2025)
A Split-Client Approach to Second-Order Optimization
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2025)
von: Chayti, El Mahdi, et al.
Veröffentlicht: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
Learning to Forget: Continual Learning with Adaptive Weight Decay
von: Ramesh, Aditya A., et al.
Veröffentlicht: (2026)
von: Ramesh, Aditya A., et al.
Veröffentlicht: (2026)
Why Do We Need Weight Decay in Modern Deep Learning?
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
von: D'Angelo, Francesco, et al.
Veröffentlicht: (2023)
The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression
von: Xu, Yongzhong
Veröffentlicht: (2026)
von: Xu, Yongzhong
Veröffentlicht: (2026)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
von: Fan, Simin, et al.
Veröffentlicht: (2025)
von: Fan, Simin, et al.
Veröffentlicht: (2025)
Randomly Weighted Neuromodulation in Neural Networks Facilitates Learning of Manifolds Common Across Tasks
von: Hong, Jinyung, et al.
Veröffentlicht: (2023)
von: Hong, Jinyung, et al.
Veröffentlicht: (2023)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm Perspective
von: Xie, Zeke, et al.
Veröffentlicht: (2020)
von: Xie, Zeke, et al.
Veröffentlicht: (2020)
Bridging Equilibrium and Kinetics Prediction with a Data-Weighted Neural Network Model of Methane Steam Reforming
von: Pizoń, Zofia, et al.
Veröffentlicht: (2025)
von: Pizoń, Zofia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
von: Kosson, Atli, et al.
Veröffentlicht: (2024) -
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
von: Kosson, Atli, et al.
Veröffentlicht: (2025) -
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
von: Dremov, Aleksandr, et al.
Veröffentlicht: (2025) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025) -
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)