Sparse maximal update parameterization: A holistic approach to sparse training dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Dey, Nolan, Bergsma, Shane, Hestness, Joel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025)
by: Bergsma, Shane, et al.
Published: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)
by: Gray, Gavia, et al.
Published: (2024)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
MediSwift: Efficient Sparse Pre-trained Biomedical Language Models
by: Thangarasa, Vithursan, et al.
Published: (2024)
by: Thangarasa, Vithursan, et al.
Published: (2024)
Efficiently Disentangle Causal Representations
by: Li, Yuanpeng, et al.
Published: (2022)
by: Li, Yuanpeng, et al.
Published: (2022)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
CoRPO: Adding a Correctness Bias to GRPO Improves Generalization
by: Garg, Anisha, et al.
Published: (2025)
by: Garg, Anisha, et al.
Published: (2025)
Reconstructing dynamics from sparse observations with no training on target system
by: Zhai, Zheng-Meng, et al.
Published: (2024)
by: Zhai, Zheng-Meng, et al.
Published: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
A joint optimization approach to identifying sparse dynamics using least squares kernel collocation
by: Hsu, Alexander W., et al.
Published: (2025)
by: Hsu, Alexander W., et al.
Published: (2025)
Generative Pre-trained Ranking Model with Over-parameterization at Web-Scale (Extended Abstract)
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Sparse deep neural networks for nonparametric estimation in high-dimensional sparse regression
by: Wu, Dongya, et al.
Published: (2024)
by: Wu, Dongya, et al.
Published: (2024)
Robustness in sparse artificial neural networks trained with adaptive topology
by: Sulyok, Bendegúz, et al.
Published: (2026)
by: Sulyok, Bendegúz, et al.
Published: (2026)
Random Sparse Lifts: Construction, Analysis and Convergence of finite sparse networks
by: Robin, David A. R., et al.
Published: (2025)
by: Robin, David A. R., et al.
Published: (2025)
On-line learning of dynamic systems: sparse regression meets Kalman filtering
by: Pillonetto, Gianluigi, et al.
Published: (2025)
by: Pillonetto, Gianluigi, et al.
Published: (2025)
Multiplicative update rules for accelerating deep learning training and increasing robustness
by: Kirtas, Manos, et al.
Published: (2023)
by: Kirtas, Manos, et al.
Published: (2023)
A sparse PAC-Bayesian approach for high-dimensional quantile prediction
by: Mai, The Tien
Published: (2024)
by: Mai, The Tien
Published: (2024)
On the Benefits of Over-parameterization for Out-of-Distribution Generalization
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
One-for-All: A Lightweight Stabilized and Parameter-Efficient Pre-trained LLM for Time Series Forecasting
by: Dey, Prasanjit, et al.
Published: (2026)
by: Dey, Prasanjit, et al.
Published: (2026)
SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Pre-Defined Sparse Neural Networks with Hardware Acceleration
by: Dey, Sourya, et al.
Published: (2018)
by: Dey, Sourya, et al.
Published: (2018)
LOST: Low-rank and Sparse Pre-training for Large Language Models
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Detecting underdetermination in parameterized quantum circuits
by: Kempkes, Marie, et al.
Published: (2025)
by: Kempkes, Marie, et al.
Published: (2025)
A semi-supervised learning using over-parameterized regression
by: Hagiwara, Katsuyuki
Published: (2024)
by: Hagiwara, Katsuyuki
Published: (2024)
Intrinsic training dynamics of deep neural networks
by: Marcotte, Sibylle, et al.
Published: (2025)
by: Marcotte, Sibylle, et al.
Published: (2025)
S-STE: Continuous Pruning Function for Efficient 2:4 Sparse Pre-training
by: Hu, Yuezhou, et al.
Published: (2024)
by: Hu, Yuezhou, et al.
Published: (2024)
Mesh-free sparse identification of nonlinear dynamics
by: Gao, Mars Liyao, et al.
Published: (2025)
by: Gao, Mars Liyao, et al.
Published: (2025)
Optimal sparse phase retrieval via a quasi-Bayesian approach
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
Towards a holistic understanding of Selection Bias for Causal Effect Identification
by: Qiu, Yiwen, et al.
Published: (2026)
by: Qiu, Yiwen, et al.
Published: (2026)
sparseGeoHOPCA: A Geometric Solution to Sparse Higher-Order PCA Without Covariance Estimation
by: Xu, Renjie, et al.
Published: (2025)
by: Xu, Renjie, et al.
Published: (2025)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
by: Liu, Fanghui, et al.
Published: (2024)
by: Liu, Fanghui, et al.
Published: (2024)
An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
by: Li, Guanchen, et al.
Published: (2024)
by: Li, Guanchen, et al.
Published: (2024)
Stochastic Trust-Region Methods for Over-parameterized Models
by: Yang, Aike, et al.
Published: (2026)
by: Yang, Aike, et al.
Published: (2026)
How to find expressible and trainable parameterized quantum circuits?
by: Röseler, Peter, et al.
Published: (2026)
by: Röseler, Peter, et al.
Published: (2026)
On the Convergence of Federated Averaging under Partial Participation for Over-parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2023)
by: Liu, Xin, et al.
Published: (2023)
Similar Items
-
Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training
by: Bergsma, Shane, et al.
Published: (2025) -
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
by: Bergsma, Shane, et al.
Published: (2025) -
Scaling with Collapse: Efficient and Predictable Training of LLM Families
by: Bergsma, Shane, et al.
Published: (2025) -
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)