Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
Fuente:
arXiv
Guardado en:
| Autores principales: | Qiu, Shikai, Chen, Zixi, Phan, Hoang, Lei, Qi, Wilson, Andrew Gordon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Controllable Prompt Tuning For Balancing Group Distributional Robustness
por: Phan, Hoang, et al.
Publicado: (2024)
por: Phan, Hoang, et al.
Publicado: (2024)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
por: Qiu, Shikai, et al.
Publicado: (2025)
por: Qiu, Shikai, et al.
Publicado: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
por: Gruver, Nate, et al.
Publicado: (2023)
por: Gruver, Nate, et al.
Publicado: (2023)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
por: Qiu, Shikai, et al.
Publicado: (2024)
por: Qiu, Shikai, et al.
Publicado: (2024)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)
por: Marek, Martin, et al.
Publicado: (2026)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
por: Qiu, Shikai, et al.
Publicado: (2024)
por: Qiu, Shikai, et al.
Publicado: (2024)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
por: Potapczynski, Andres, et al.
Publicado: (2024)
por: Potapczynski, Andres, et al.
Publicado: (2024)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
por: Kuang, Yilun, et al.
Publicado: (2025)
por: Kuang, Yilun, et al.
Publicado: (2025)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
por: Finzi, Marc, et al.
Publicado: (2026)
por: Finzi, Marc, et al.
Publicado: (2026)
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
por: Xu, Ruihan, et al.
Publicado: (2026)
por: Xu, Ruihan, et al.
Publicado: (2026)
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
por: Deng, Shenyang, et al.
Publicado: (2026)
por: Deng, Shenyang, et al.
Publicado: (2026)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2025)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2025)
Elucidating the Preconditioning in Consistency Distillation
por: Zheng, Kaiwen, et al.
Publicado: (2025)
por: Zheng, Kaiwen, et al.
Publicado: (2025)
Principled Architecture-aware Scaling of Hyperparameters
por: Chen, Wuyang, et al.
Publicado: (2024)
por: Chen, Wuyang, et al.
Publicado: (2024)
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
por: Shulgin, Egor, et al.
Publicado: (2026)
por: Shulgin, Egor, et al.
Publicado: (2026)
Hyperparameter Transfer for Dense Associative Memories
por: Holtzman, Roi, et al.
Publicado: (2026)
por: Holtzman, Roi, et al.
Publicado: (2026)
Hyperparameter Transfer with Mixture-of-Expert Layers
por: Jiang, Tianze, et al.
Publicado: (2026)
por: Jiang, Tianze, et al.
Publicado: (2026)
Overtuning in Hyperparameter Optimization
por: Schneider, Lennart, et al.
Publicado: (2025)
por: Schneider, Lennart, et al.
Publicado: (2025)
Deep Learning is Not So Mysterious or Different
por: Wilson, Andrew Gordon
Publicado: (2025)
por: Wilson, Andrew Gordon
Publicado: (2025)
Out-of-Distribution Detection Methods Answer the Wrong Questions
por: Li, Yucen Lily, et al.
Publicado: (2025)
por: Li, Yucen Lily, et al.
Publicado: (2025)
Sketchy Moment Matching: Toward Fast and Provable Data Selection for Finetuning
por: Dong, Yijun, et al.
Publicado: (2024)
por: Dong, Yijun, et al.
Publicado: (2024)
Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional Responses
por: Liu, Hui, et al.
Publicado: (2024)
por: Liu, Hui, et al.
Publicado: (2024)
Hyperparameter Optimization in Machine Learning
por: Franceschi, Luca, et al.
Publicado: (2024)
por: Franceschi, Luca, et al.
Publicado: (2024)
Understanding the Mechanisms of Fast Hyperparameter Transfer
por: Ghosh, Nikhil, et al.
Publicado: (2025)
por: Ghosh, Nikhil, et al.
Publicado: (2025)
Dynamic Priors in Bayesian Optimization for Hyperparameter Optimization
por: Fehring, Lukas, et al.
Publicado: (2025)
por: Fehring, Lukas, et al.
Publicado: (2025)
Toward a Holistic Approach to Continual Model Merging
por: Phan, Hoang, et al.
Publicado: (2025)
por: Phan, Hoang, et al.
Publicado: (2025)
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
por: Phan, Hoang, et al.
Publicado: (2025)
por: Phan, Hoang, et al.
Publicado: (2025)
On Optimal Hyperparameters for Differentially Private Deep Transfer Learning
por: Rehn, Aki, et al.
Publicado: (2025)
por: Rehn, Aki, et al.
Publicado: (2025)
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks
por: Wu, Shenxi, et al.
Publicado: (2026)
por: Wu, Shenxi, et al.
Publicado: (2026)
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization
por: Rakotoarison, Herilalaina, et al.
Publicado: (2024)
por: Rakotoarison, Herilalaina, et al.
Publicado: (2024)
Exploring the Optimized Value of Each Hyperparameter in Various Gradient Descent Algorithms
por: Chen, Abel C. H.
Publicado: (2022)
por: Chen, Abel C. H.
Publicado: (2022)
Enhancing Performance and Calibration in Quantile Hyperparameter Optimization
por: Doyle, Riccardo
Publicado: (2025)
por: Doyle, Riccardo
Publicado: (2025)
Adaptive Hyperparameter Optimization for Continual Learning Scenarios
por: Semola, Rudy, et al.
Publicado: (2024)
por: Semola, Rudy, et al.
Publicado: (2024)
Hyperparameter Tuning Through Pessimistic Bilevel Optimization
por: Ustun, Meltem Apaydin, et al.
Publicado: (2024)
por: Ustun, Meltem Apaydin, et al.
Publicado: (2024)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
por: Zhang, Gavin, et al.
Publicado: (2025)
por: Zhang, Gavin, et al.
Publicado: (2025)
CASHomon Sets: Efficient Rashomon Sets Across Multiple Model Classes and their Hyperparameters
por: Ewald, Fiona Katharina, et al.
Publicado: (2026)
por: Ewald, Fiona Katharina, et al.
Publicado: (2026)
LLM Circuit Analyses Are Consistent Across Training and Scale
por: Tigges, Curt, et al.
Publicado: (2024)
por: Tigges, Curt, et al.
Publicado: (2024)
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
por: Khaddaj, Alaa, et al.
Publicado: (2025)
por: Khaddaj, Alaa, et al.
Publicado: (2025)
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix
por: Yue, Yun, et al.
Publicado: (2023)
por: Yue, Yun, et al.
Publicado: (2023)
Scaling Exponents Across Parameterizations and Optimizers
por: Everett, Katie, et al.
Publicado: (2024)
por: Everett, Katie, et al.
Publicado: (2024)
Ejemplares similares
-
Controllable Prompt Tuning For Balancing Group Distributional Robustness
por: Phan, Hoang, et al.
Publicado: (2024) -
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
por: Qiu, Shikai, et al.
Publicado: (2025) -
Large Language Models Are Zero-Shot Time Series Forecasters
por: Gruver, Nate, et al.
Publicado: (2023) -
Transferring Knowledge from Large Foundation Models to Small Downstream Models
por: Qiu, Shikai, et al.
Publicado: (2024) -
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)