The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Tian, Humayun, Ahmed Imtiaz, Evci, Utku, Subramanian, Suvinay, Yazdanbakhsh, Amir, Alistarh, Dan, Dziugaite, Gintare Karolina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024)
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
by: Ohib, Riyasat, et al.
Published: (2024)
by: Ohib, Riyasat, et al.
Published: (2024)
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
by: Üyük, Cem, et al.
Published: (2024)
by: Üyük, Cem, et al.
Published: (2024)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
by: Jiang, Wenqi, et al.
Published: (2025)
by: Jiang, Wenqi, et al.
Published: (2025)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
by: Yang, Yu, et al.
Published: (2023)
by: Yang, Yu, et al.
Published: (2023)
Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design
by: Yazdanbakhsh, Amir
Published: (2025)
by: Yazdanbakhsh, Amir
Published: (2025)
Data Selection for Transfer Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2024)
Dynamic Sparse Training with Structured Sparsity
by: Lasby, Mike, et al.
Published: (2023)
by: Lasby, Mike, et al.
Published: (2023)
Leveraging Function Space Aggregation for Federated Learning at Scale
by: Dhawan, Nikita, et al.
Published: (2023)
by: Dhawan, Nikita, et al.
Published: (2023)
Dataset Difficulty and the Role of Inductive Bias
by: Kwok, Devin, et al.
Published: (2024)
by: Kwok, Devin, et al.
Published: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
by: Guo, Phillip, et al.
Published: (2024)
by: Guo, Phillip, et al.
Published: (2024)
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
by: Baluta, Teodora, et al.
Published: (2024)
by: Baluta, Teodora, et al.
Published: (2024)
Information Complexity of Stochastic Convex Optimization: Applications to Generalization and Memorization
by: Attias, Idan, et al.
Published: (2024)
by: Attias, Idan, et al.
Published: (2024)
Improved Localized Machine Unlearning Through the Lens of Memorization
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
by: Torkzadehmahani, Reihaneh, et al.
Published: (2024)
Expand Neurons, Not Parameters
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Simultaneous linear connectivity of neural networks modulo permutation
by: Sharma, Ekansh, et al.
Published: (2024)
by: Sharma, Ekansh, et al.
Published: (2024)
On Traceability in $\ell_p$ Stochastic Convex Optimization
by: Voitovych, Sasha, et al.
Published: (2025)
by: Voitovych, Sasha, et al.
Published: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
by: Vishwanathan, Manoj, et al.
Published: (2026)
by: Vishwanathan, Manoj, et al.
Published: (2026)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
Detoxifying LLMs via Representation Erasure-Based Preference Optimization
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2026)
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2025)
Torque-Aware Momentum
by: Malviya, Pranshu, et al.
Published: (2024)
by: Malviya, Pranshu, et al.
Published: (2024)
The Geometric Structure of Models Learning Sparse Data
by: Walker, Thomas, et al.
Published: (2026)
by: Walker, Thomas, et al.
Published: (2026)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
Effective Interplay between Sparsity and Quantization: From Theory to Practice
by: Harma, Simla Burcu, et al.
Published: (2024)
by: Harma, Simla Burcu, et al.
Published: (2024)
Towards Robust Scaling Laws for Optimizers
by: Volkova, Alexandra, et al.
Published: (2026)
by: Volkova, Alexandra, et al.
Published: (2026)
Mitigating over-exploration in latent space optimization using LES
by: Ronen, Omer, et al.
Published: (2024)
by: Ronen, Omer, et al.
Published: (2024)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
by: Li, Guanchen, et al.
Published: (2024)
by: Li, Guanchen, et al.
Published: (2024)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
Leveraging Per-Instance Privacy for Machine Unlearning
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
by: Sepahvand, Nazanin Mohammadi, et al.
Published: (2025)
Similar Items
-
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2024) -
SSFL: Discovering Sparse Unified Subnetworks at Initialization for Efficient Federated Learning
by: Ohib, Riyasat, et al.
Published: (2024) -
The Non-Local Model Merging Problem: Permutation Symmetries and Variance Collapse
by: Sharma, Ekansh, et al.
Published: (2024) -
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
by: Üyük, Cem, et al.
Published: (2024)