P$^2$ Law: Scaling Law for Post-Training After Model Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909622759063552 |
|---|---|
| author | Chen, Xiaodong Hu, Yuxuan Zhang, Xiaokang Wang, Yanling Li, Cuiping Chen, Hong Zhang, Jing |
| author_facet | Chen, Xiaodong Hu, Yuxuan Zhang, Xiaokang Wang, Yanling Li, Cuiping Chen, Hong Zhang, Jing |
| contents | Pruning has become a widely adopted technique for reducing the hardware requirements of large language models (LLMs). To recover model performance after pruning, post-training is commonly employed to mitigate the resulting performance degradation. While post-training benefits from larger datasets, once the dataset size is already substantial, increasing the training data provides only limited performance gains. To balance post-training cost and model performance, it is necessary to explore the optimal amount of post-training data.Through extensive experiments on the Llama-3 and Qwen-2.5 series models, pruned using various common pruning methods, we uncover the scaling \textbf{Law} for \textbf{P}ost-training after model \textbf{P}runing, referred to as the P$^2$ Law.This law identifies four key factors for predicting the pruned model's post-training loss: the model size before pruning, the number of post-training tokens, the pruning rate, and the model's loss before pruning. Moreover, P$^2$ Law can generalize to larger dataset sizes, larger model sizes, and higher pruning rates, offering valuable insights for the post-training of pruned LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_10272 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | P$^2$ Law: Scaling Law for Post-Training After Model Pruning Chen, Xiaodong Hu, Yuxuan Zhang, Xiaokang Wang, Yanling Li, Cuiping Chen, Hong Zhang, Jing Artificial Intelligence Computation and Language Machine Learning Pruning has become a widely adopted technique for reducing the hardware requirements of large language models (LLMs). To recover model performance after pruning, post-training is commonly employed to mitigate the resulting performance degradation. While post-training benefits from larger datasets, once the dataset size is already substantial, increasing the training data provides only limited performance gains. To balance post-training cost and model performance, it is necessary to explore the optimal amount of post-training data.Through extensive experiments on the Llama-3 and Qwen-2.5 series models, pruned using various common pruning methods, we uncover the scaling \textbf{Law} for \textbf{P}ost-training after model \textbf{P}runing, referred to as the P$^2$ Law.This law identifies four key factors for predicting the pruned model's post-training loss: the model size before pruning, the number of post-training tokens, the pruning rate, and the model's loss before pruning. Moreover, P$^2$ Law can generalize to larger dataset sizes, larger model sizes, and higher pruning rates, offering valuable insights for the post-training of pruned LLMs. |
| title | P$^2$ Law: Scaling Law for Post-Training After Model Pruning |
| topic | Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2411.10272 |