Hyperparameter Transfer with Mixture-of-Expert Layers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Tianze, Bordelon, Blake, Pehlevan, Cengiz, Hanin, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026)
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025)
A Dynamical Model of Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Infinite Limits of Multi-head Transformer Dynamics
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Grokking as the Transition from Lazy to Rich Training Dynamics
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
Hyperparameter Transfer for Dense Associative Memories
von: Holtzman, Roi, et al.
Veröffentlicht: (2026)
von: Holtzman, Roi, et al.
Veröffentlicht: (2026)
Don't be lazy: CompleteP enables compute-efficient deep transformers
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
von: Dey, Nolan, et al.
Veröffentlicht: (2025)
Bayesian Inference with Shaped Deep Non-linear MLPs
von: Hanin, Boris, et al.
Veröffentlicht: (2026)
von: Hanin, Boris, et al.
Veröffentlicht: (2026)
Dynamically Learning to Integrate in Recurrent Neural Networks
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
von: Bordelon, Blake, et al.
Veröffentlicht: (2025)
Do Mice Grok? Glimpses of Hidden Progress During Overtraining in Sensory Cortex
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2025)
Principled Architecture-aware Scaling of Hyperparameters
von: Chen, Wuyang, et al.
Veröffentlicht: (2024)
von: Chen, Wuyang, et al.
Veröffentlicht: (2024)
Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
von: Halder, Indranil, et al.
Veröffentlicht: (2025)
Learning Rate Transfer in Normalized Transformers
von: Shigida, Boris, et al.
Veröffentlicht: (2026)
von: Shigida, Boris, et al.
Veröffentlicht: (2026)
Scaling Laws for Precision
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Summary statistics of learning link changing neural representations to behavior
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2025)
Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge Ensembles
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
von: Ruben, Benjamin S., et al.
Veröffentlicht: (2023)
Learning richness modulates equality reasoning in neural networks
von: Tong, William L., et al.
Veröffentlicht: (2025)
von: Tong, William L., et al.
Veröffentlicht: (2025)
MLPs Learn In-Context on Regression and Classification Tasks
von: Tong, William L., et al.
Veröffentlicht: (2024)
von: Tong, William L., et al.
Veröffentlicht: (2024)
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
von: Halder, Indranil, et al.
Veröffentlicht: (2026)
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
Convex Relaxation for Solving Large-Margin Classifiers in Hyperbolic Space
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
von: Yang, Sheng, et al.
Veröffentlicht: (2024)
Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
von: Bordelon, Blake, et al.
Veröffentlicht: (2026)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Nadaraya-Watson kernel smoothing as a random energy model
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
von: Zavatone-Veth, Jacob A., et al.
Veröffentlicht: (2024)
Correlative Information Maximization: A Biologically Plausible Approach to Supervised Deep Neural Networks without Weight Symmetry
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2023)
von: Bozkurt, Bariscan, et al.
Veröffentlicht: (2023)
A solvable model of learning generative diffusion: theory and insights
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
von: Cui, Hugo, et al.
Veröffentlicht: (2025)
The Optimization Landscape of SGD Across the Feature Learning Strength
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
von: Uzun, Mustafa, et al.
Veröffentlicht: (2026)
von: Uzun, Mustafa, et al.
Veröffentlicht: (2026)
Universal One-third Time Scaling in Learning Peaked Distributions
von: Liu, Yizhou, et al.
Veröffentlicht: (2026)
von: Liu, Yizhou, et al.
Veröffentlicht: (2026)
Pixel-Based Similarities as an Alternative to Neural Data for Improving Convolutional Neural Network Adversarial Robustness
von: Attias, Elie, et al.
Veröffentlicht: (2024)
von: Attias, Elie, et al.
Veröffentlicht: (2024)
Global Universality of Singular Values in Products of Many Large Random Matrices
von: Hanin, Boris, et al.
Veröffentlicht: (2025)
von: Hanin, Boris, et al.
Veröffentlicht: (2025)
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
von: Tong, William L., et al.
Veröffentlicht: (2026)
von: Tong, William L., et al.
Veröffentlicht: (2026)
An Analytical Theory of Spectral Bias in the Learning Dynamics of Diffusion Models
von: Wang, Binxu, et al.
Veröffentlicht: (2025)
von: Wang, Binxu, et al.
Veröffentlicht: (2025)
Risk and cross validation in ridge regression with correlated samples
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
von: Atanasov, Alexander, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
von: Bordelon, Blake, et al.
Veröffentlicht: (2025) -
Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning
von: Bordelon, Blake, et al.
Veröffentlicht: (2026) -
Transfer Learning in Infinite Width Feature Learning Networks
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2025) -
Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer
von: Lauditi, Clarissa, et al.
Veröffentlicht: (2026) -
How Feature Learning Can Improve Neural Scaling Laws
von: Bordelon, Blake, et al.
Veröffentlicht: (2024)