PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Bowei, Yin, Lihao, Zhen, Hui-Ling, Zhang, Xiaokun, Yuan, Mingxuan, Ma, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911442712657920
author He, Bowei
Yin, Lihao
Zhen, Hui-Ling
Zhang, Xiaokun
Yuan, Mingxuan
Ma, Chen
author_facet He, Bowei
Yin, Lihao
Zhen, Hui-Ling
Zhang, Xiaokun
Yuan, Mingxuan
Ma, Chen
contents Model pruning is an effective approach for compressing large language models (LLMs). However, this process often leads to significant degradation of model capabilities. While post-training techniques such as instruction tuning are commonly employed to recover model performance, existing methods often overlook the uneven deterioration of model capabilities and incur high computational costs. Moreover, some irrelevant instructions may also introduce negative effects to model capacity recovery. To address these challenges, we propose the \textbf{P}ost-training d\textbf{A}ta \textbf{S}election method for \textbf{E}fficient pruned large language model \textbf{R}ecovery (\textbf{PASER}). PASER aims to identify instructions to recover the most compromised model capacities with a certain data budget. Our approach first applies manifold learning and spectral clustering to group recovery instructions in the semantic space, revealing capability-specific instruction sets. Then, the data budget is adaptively allocated across clusters by the degree of corresponding model capability degradation. In each cluster, we prioritize data samples that lead to the most decline of model performance. To mitigate potential negative tuning effects, we also detect and filter out conflicting or irrelevant recovery data. Extensive experiments demonstrate that PASER significantly outperforms conventional baselines, effectively recovering the general capabilities of pruned LLMs while utilizing merely 4\%-20\% of the original post-training data. We provide the code repository in \href{https://github.com/BokwaiHo/PASER}{Link}.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
He, Bowei
Yin, Lihao
Zhen, Hui-Ling
Zhang, Xiaokun
Yuan, Mingxuan
Ma, Chen
Computation and Language
Model pruning is an effective approach for compressing large language models (LLMs). However, this process often leads to significant degradation of model capabilities. While post-training techniques such as instruction tuning are commonly employed to recover model performance, existing methods often overlook the uneven deterioration of model capabilities and incur high computational costs. Moreover, some irrelevant instructions may also introduce negative effects to model capacity recovery. To address these challenges, we propose the \textbf{P}ost-training d\textbf{A}ta \textbf{S}election method for \textbf{E}fficient pruned large language model \textbf{R}ecovery (\textbf{PASER}). PASER aims to identify instructions to recover the most compromised model capacities with a certain data budget. Our approach first applies manifold learning and spectral clustering to group recovery instructions in the semantic space, revealing capability-specific instruction sets. Then, the data budget is adaptively allocated across clusters by the degree of corresponding model capability degradation. In each cluster, we prioritize data samples that lead to the most decline of model performance. To mitigate potential negative tuning effects, we also detect and filter out conflicting or irrelevant recovery data. Extensive experiments demonstrate that PASER significantly outperforms conventional baselines, effectively recovering the general capabilities of pruned LLMs while utilizing merely 4\%-20\% of the original post-training data. We provide the code repository in \href{https://github.com/BokwaiHo/PASER}{Link}.
title PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
topic Computation and Language
url https://arxiv.org/abs/2502.12594