Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shulgin, Egor, von Rütte, Dimitri, Zhang, Tianyue H., Ajroldi, Niccolò, Schölkopf, Bernhard, Orvieto, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Behavior of Discrete Diffusion Language Models
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
Generalized Interpolating Discrete Diffusion
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
von: Riabinin, Artem, et al.
Veröffentlicht: (2025)
von: Riabinin, Artem, et al.
Veröffentlicht: (2025)
On the Convergence of DP-SGD with Adaptive Clipping
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
von: Shulgin, Egor, et al.
Veröffentlicht: (2024)
Towards a Better Theoretical Understanding of Independent Subnetwork Training
von: Shulgin, Egor, et al.
Veröffentlicht: (2023)
von: Shulgin, Egor, et al.
Veröffentlicht: (2023)
Training Dynamics Impact Post-Training Quantization Robustness
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
Smoothed Normalization for Efficient Distributed Private Optimization
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
FIGARO: Generating Symbolic Music with Fine-Grained Artistic Control
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2022)
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2022)
First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
von: Bill, Eric Tillman, et al.
Veröffentlicht: (2025)
von: Bill, Eric Tillman, et al.
Veröffentlicht: (2025)
Beyond the Ideal: Analyzing the Inexact Muon Update
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
von: Shulgin, Egor, et al.
Veröffentlicht: (2025)
Skill or Luck? Return Decomposition via Advantage Functions
von: Pan, Hsiao-Ru, et al.
Veröffentlicht: (2024)
von: Pan, Hsiao-Ru, et al.
Veröffentlicht: (2024)
SPAM: Stochastic Proximal Point Method with Momentum Variance Reduction for Non-convex Cross-Device Federated Learning
von: Karagulyan, Avetik, et al.
Veröffentlicht: (2024)
von: Karagulyan, Avetik, et al.
Veröffentlicht: (2024)
Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
von: Da Costa, Lancelot, et al.
Veröffentlicht: (2025)
von: Da Costa, Lancelot, et al.
Veröffentlicht: (2025)
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025)
von: Laing, Sam, et al.
Veröffentlicht: (2025)
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
Geometry-Aware Instrumental Variable Regression
von: Kremer, Heiner, et al.
Veröffentlicht: (2024)
von: Kremer, Heiner, et al.
Veröffentlicht: (2024)
Robustness of Nonlinear Representation Learning
von: Buchholz, Simon, et al.
Veröffentlicht: (2025)
von: Buchholz, Simon, et al.
Veröffentlicht: (2025)
MAST: Model-Agnostic Sparsified Training
von: Demidovich, Yury, et al.
Veröffentlicht: (2023)
von: Demidovich, Yury, et al.
Veröffentlicht: (2023)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
In Search of Adam's Secret Sauce
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
Overtuning in Hyperparameter Optimization
von: Schneider, Lennart, et al.
Veröffentlicht: (2025)
von: Schneider, Lennart, et al.
Veröffentlicht: (2025)
Physics of Learning: A Lagrangian perspective to different learning paradigms
von: Guo, Siyuan, et al.
Veröffentlicht: (2025)
von: Guo, Siyuan, et al.
Veröffentlicht: (2025)
crowd-hpo: Realistic Hyperparameter Optimization and Benchmarking for Learning from Crowds with Noisy Labels
von: Herde, Marek, et al.
Veröffentlicht: (2025)
von: Herde, Marek, et al.
Veröffentlicht: (2025)
Muown: Row-Norm Control for Muon Optimization
von: Lion, Kai, et al.
Veröffentlicht: (2026)
von: Lion, Kai, et al.
Veröffentlicht: (2026)
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks
von: Wu, Shenxi, et al.
Veröffentlicht: (2026)
von: Wu, Shenxi, et al.
Veröffentlicht: (2026)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
Explaining Grokking in Transformers through the Lens of Inductive Bias
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
Improved state mixing in higher-order and block diagonal linear recurrent networks
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
von: Dubinin, Igor, et al.
Veröffentlicht: (2026)
An Uncertainty Principle for Linear Recurrent Neural Networks
von: François, Alexandre, et al.
Veröffentlicht: (2025)
von: François, Alexandre, et al.
Veröffentlicht: (2025)
Causal Modeling with Stationary Diffusions
von: Lorch, Lars, et al.
Veröffentlicht: (2023)
von: Lorch, Lars, et al.
Veröffentlicht: (2023)
SPARTAN: A Sparse Transformer World Model Attending to What Matters
von: Lei, Anson, et al.
Veröffentlicht: (2024)
von: Lei, Anson, et al.
Veröffentlicht: (2024)
Online Learning and Unlearning
von: Hu, Yaxi, et al.
Veröffentlicht: (2025)
von: Hu, Yaxi, et al.
Veröffentlicht: (2025)
Targeted Reduction of Causal Models
von: Kekić, Armin, et al.
Veröffentlicht: (2023)
von: Kekić, Armin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Scaling Behavior of Discrete Diffusion Language Models
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025) -
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025) -
Generalized Interpolating Discrete Diffusion
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025) -
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025) -
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)