Saved in:
| Main Authors: | Kosson, Atli, Messmer, Bettina, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.23922 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
by: Kosson, Atli, et al.
Published: (2025)
by: Kosson, Atli, et al.
Published: (2025)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024)
by: Hägele, Alexander, et al.
Published: (2024)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
by: Kalra, Dayal Singh, et al.
Published: (2024)
by: Kalra, Dayal Singh, et al.
Published: (2024)
Taming Transformer Without Using Learning Rate Warmup
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
by: Liu, Yuxing, et al.
Published: (2025)
by: Liu, Yuxing, et al.
Published: (2025)
Efficient Generative Model Training via Embedded Representation Warmup
by: Liu, Deyuan, et al.
Published: (2025)
by: Liu, Deyuan, et al.
Published: (2025)
Data Warmup: Complexity-Aware Curricula for Efficient Diffusion Training
by: Lin, Jinhong, et al.
Published: (2026)
by: Lin, Jinhong, et al.
Published: (2026)
Optimal Learning-Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay
by: Li, Binghui, et al.
Published: (2026)
by: Li, Binghui, et al.
Published: (2026)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Unified Convergence Theory of Stochastic and Variance-Reduced Cubic Newton Methods
by: Chayti, El Mahdi, et al.
Published: (2023)
by: Chayti, El Mahdi, et al.
Published: (2023)
FedPeWS: Personalized Warmup via Subnetworks for Enhanced Heterogeneous Federated Learning
by: Tastan, Nurbek, et al.
Published: (2024)
by: Tastan, Nurbek, et al.
Published: (2024)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
A New First-Order Meta-Learning Algorithm with Convergence Guarantees
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
Test-Time Warmup for Multimodal Large Language Models
by: Rajaneesh, Nikita, et al.
Published: (2025)
by: Rajaneesh, Nikita, et al.
Published: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
by: Belloni, Annalisa, et al.
Published: (2026)
by: Belloni, Annalisa, et al.
Published: (2026)
CoBo: Collaborative Learning via Bilevel Optimization
by: Hashemi, Diba, et al.
Published: (2024)
by: Hashemi, Diba, et al.
Published: (2024)
Using Machine Learning for move sequence visualization and generation in climbing
by: Rimbot, Thomas, et al.
Published: (2025)
by: Rimbot, Thomas, et al.
Published: (2025)
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
Stochastic Difference-of-Convex Optimization with Momentum
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
A Split-Client Approach to Second-Order Optimization
by: Chayti, El Mahdi, et al.
Published: (2025)
by: Chayti, El Mahdi, et al.
Published: (2025)
Accuracy Booster: Enabling 4-bit Fixed-point Arithmetic for DNN Training
by: Harma, Simla Burcu, et al.
Published: (2022)
by: Harma, Simla Burcu, et al.
Published: (2022)
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
by: Fan, Simin, et al.
Published: (2025)
by: Fan, Simin, et al.
Published: (2025)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Local MixVR: Breaking the Communication-Sample Dependence in Distributed Learning
by: Dahan, Tehila, et al.
Published: (2026)
by: Dahan, Tehila, et al.
Published: (2026)
Apertus LLM Family Expansion via Distillation and Quantization
by: Panferov, Andrei, et al.
Published: (2026)
by: Panferov, Andrei, et al.
Published: (2026)
Improving Stochastic Cubic Newton with Momentum
by: Chayti, El Mahdi, et al.
Published: (2024)
by: Chayti, El Mahdi, et al.
Published: (2024)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Combining Abstract Argumentation and Machine Learning for Efficiently Analyzing Low-Level Process Event Streams
by: Fazzinga, Bettina, et al.
Published: (2025)
by: Fazzinga, Bettina, et al.
Published: (2025)
Similar Items
-
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023) -
Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
by: Dremov, Aleksandr, et al.
Published: (2025) -
Weight Decay may matter more than muP for Learning Rate Transfer in Practice
by: Kosson, Atli, et al.
Published: (2025) -
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)