P$^2$ Law: Scaling Law for Post-Training After Model Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiaodong, Hu, Yuxuan, Zhang, Xiaokang, Wang, Yanling, Li, Cuiping, Chen, Hong, Zhang, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909622759063552
author Chen, Xiaodong
Hu, Yuxuan
Zhang, Xiaokang
Wang, Yanling
Li, Cuiping
Chen, Hong
Zhang, Jing
author_facet Chen, Xiaodong
Hu, Yuxuan
Zhang, Xiaokang
Wang, Yanling
Li, Cuiping
Chen, Hong
Zhang, Jing
contents Pruning has become a widely adopted technique for reducing the hardware requirements of large language models (LLMs). To recover model performance after pruning, post-training is commonly employed to mitigate the resulting performance degradation. While post-training benefits from larger datasets, once the dataset size is already substantial, increasing the training data provides only limited performance gains. To balance post-training cost and model performance, it is necessary to explore the optimal amount of post-training data.Through extensive experiments on the Llama-3 and Qwen-2.5 series models, pruned using various common pruning methods, we uncover the scaling \textbf{Law} for \textbf{P}ost-training after model \textbf{P}runing, referred to as the P$^2$ Law.This law identifies four key factors for predicting the pruned model's post-training loss: the model size before pruning, the number of post-training tokens, the pruning rate, and the model's loss before pruning. Moreover, P$^2$ Law can generalize to larger dataset sizes, larger model sizes, and higher pruning rates, offering valuable insights for the post-training of pruned LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle P$^2$ Law: Scaling Law for Post-Training After Model Pruning
Chen, Xiaodong
Hu, Yuxuan
Zhang, Xiaokang
Wang, Yanling
Li, Cuiping
Chen, Hong
Zhang, Jing
Artificial Intelligence
Computation and Language
Machine Learning
Pruning has become a widely adopted technique for reducing the hardware requirements of large language models (LLMs). To recover model performance after pruning, post-training is commonly employed to mitigate the resulting performance degradation. While post-training benefits from larger datasets, once the dataset size is already substantial, increasing the training data provides only limited performance gains. To balance post-training cost and model performance, it is necessary to explore the optimal amount of post-training data.Through extensive experiments on the Llama-3 and Qwen-2.5 series models, pruned using various common pruning methods, we uncover the scaling \textbf{Law} for \textbf{P}ost-training after model \textbf{P}runing, referred to as the P$^2$ Law.This law identifies four key factors for predicting the pruned model's post-training loss: the model size before pruning, the number of post-training tokens, the pruning rate, and the model's loss before pruning. Moreover, P$^2$ Law can generalize to larger dataset sizes, larger model sizes, and higher pruning rates, offering valuable insights for the post-training of pruned LLMs.
title P$^2$ Law: Scaling Law for Post-Training After Model Pruning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.10272