$μ$pscaling small models: Principled warm starts and hyperparameter transfer
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Yuxin, Chen, Nan, Díaz, Mateo, Hayou, Soufiane, Kunisky, Dmitriy, Villar, Soledad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
by: Chen, Nan, et al.
Published: (2026)
by: Chen, Nan, et al.
Published: (2026)
A Proof of Learning Rate Transfer under $μ$P
by: Hayou, Soufiane
Published: (2025)
by: Hayou, Soufiane
Published: (2025)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
by: Hayou, Soufiane, et al.
Published: (2025)
by: Hayou, Soufiane, et al.
Published: (2025)
The Impact of Initialization on LoRA Finetuning Dynamics
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
LoRA+: Efficient Low Rank Adaptation of Large Models
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
Nonlinear Laplacians: Tunable principal component analysis under directional prior information
by: Ma, Yuxin, et al.
Published: (2025)
by: Ma, Yuxin, et al.
Published: (2025)
Be aware of overfitting by hyperparameter optimization!
by: Tetko, Igor V., et al.
Published: (2024)
by: Tetko, Igor V., et al.
Published: (2024)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
by: Firdoussi, Aymane El, et al.
Published: (2024)
by: Firdoussi, Aymane El, et al.
Published: (2024)
Low coordinate degree algorithms II: Categorical signals and generalized stochastic block models
by: Kunisky, Dmitriy
Published: (2024)
by: Kunisky, Dmitriy
Published: (2024)
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Auto Researching, not hyperparameter tuning: Convergence Analysis of 10,000 Experiments
by: Li, Xiaoyi
Published: (2026)
by: Li, Xiaoyi
Published: (2026)
Low coordinate degree algorithms I: Universality of computational thresholds for hypothesis testing
by: Kunisky, Dmitriy
Published: (2024)
by: Kunisky, Dmitriy
Published: (2024)
PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
by: Hayou, Soufiane, et al.
Published: (2025)
by: Hayou, Soufiane, et al.
Published: (2025)
Fuzzy hyperparameters update in a second order optimization
by: Bensadok, Abdelaziz, et al.
Published: (2024)
by: Bensadok, Abdelaziz, et al.
Published: (2024)
On Transferring Transferability: Towards a Theory for Size Generalization
by: Levin, Eitan, et al.
Published: (2025)
by: Levin, Eitan, et al.
Published: (2025)
Graph neural networks and non-commuting operators
by: Velasco, Mauricio, et al.
Published: (2024)
by: Velasco, Mauricio, et al.
Published: (2024)
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
by: Medapati, Sourabh, et al.
Published: (2025)
by: Medapati, Sourabh, et al.
Published: (2025)
An algorithmic framework for the optimization of deep neural networks architectures and hyperparameters
by: Keisler, Julie, et al.
Published: (2023)
by: Keisler, Julie, et al.
Published: (2023)
IMOVNO+: A Regional Partitioning and Meta-Heuristic Ensemble Framework for Imbalanced Multi-Class Learning
by: Bacha, Soufiane, et al.
Published: (2026)
by: Bacha, Soufiane, et al.
Published: (2026)
Leave-one-out Distinguishability in Machine Learning
by: Ye, Jiayuan, et al.
Published: (2023)
by: Ye, Jiayuan, et al.
Published: (2023)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
Tensor learning with orthogonal, Lorentz, and symplectic symmetries
by: Gregory, Wilson G., et al.
Published: (2024)
by: Gregory, Wilson G., et al.
Published: (2024)
On the transferability of Sparse Autoencoders for interpreting compressed models
by: Gupte, Suchit, et al.
Published: (2025)
by: Gupte, Suchit, et al.
Published: (2025)
Computational and statistical lower bounds for low-rank estimation under general inhomogeneous noise
by: De, Debsurya, et al.
Published: (2025)
by: De, Debsurya, et al.
Published: (2025)
State Space Models on Temporal Graphs: A First-Principles Study
by: Li, Jintang, et al.
Published: (2024)
by: Li, Jintang, et al.
Published: (2024)
Occam's model: Selecting simpler representations for better transferability estimation
by: Singh, Prabhant, et al.
Published: (2025)
by: Singh, Prabhant, et al.
Published: (2025)
An advantage based policy transfer algorithm for reinforcement learning with measures of transferability
by: Alam, Md Ferdous, et al.
Published: (2023)
by: Alam, Md Ferdous, et al.
Published: (2023)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
Principle-Evolvable Scientific Discovery via Uncertainty Minimization
by: Pu, Yingming, et al.
Published: (2026)
by: Pu, Yingming, et al.
Published: (2026)
Neural Velocity for hyperparameter tuning
by: Dalmasso, Gianluca, et al.
Published: (2025)
by: Dalmasso, Gianluca, et al.
Published: (2025)
Entropic Causal Inference: Graph Identifiability
by: Compton, Spencer, et al.
Published: (2025)
by: Compton, Spencer, et al.
Published: (2025)
DDGAD: Trajectory Dynamics for Diffusion-Based Graph Anomaly Detection
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
ANO: A Principled Approach to Robust Policy Optimization
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Expand and Compress: Exploring Tuning Principles for Continual Spatio-Temporal Graph Forecasting
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
PiFlow: Principle-Aware Scientific Discovery with Multi-Agent Collaboration
by: Pu, Yingming, et al.
Published: (2025)
by: Pu, Yingming, et al.
Published: (2025)
A Novel Double Pruning method for Imbalanced Data using Information Entropy and Roulette Wheel Selection for Breast Cancer Diagnosis
by: Bacha, Soufiane, et al.
Published: (2025)
by: Bacha, Soufiane, et al.
Published: (2025)
Leveraging Invariant Principle for Heterophilic Graph Structure Distribution Shifts
by: Yang, Jinluan, et al.
Published: (2024)
by: Yang, Jinluan, et al.
Published: (2024)
Towards Principled Graph Transformers
by: Müller, Luis, et al.
Published: (2024)
by: Müller, Luis, et al.
Published: (2024)
Preconditioning Benefits of Spectral Orthogonalization in Muon
by: Ma, Jianhao, et al.
Published: (2026)
by: Ma, Jianhao, et al.
Published: (2026)
Similar Items
-
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
by: Chen, Nan, et al.
Published: (2026) -
A Proof of Learning Rate Transfer under $μ$P
by: Hayou, Soufiane
Published: (2025) -
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
by: Hayou, Soufiane, et al.
Published: (2025) -
The Impact of Initialization on LoRA Finetuning Dynamics
by: Hayou, Soufiane, et al.
Published: (2024) -
LoRA+: Efficient Low Rank Adaptation of Large Models
by: Hayou, Soufiane, et al.
Published: (2024)