Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Yang, Ji, Yixin, Li, Juntao, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915634927894528
author Xiang, Yang
Ji, Yixin
Li, Juntao
Zhang, Min
author_facet Xiang, Yang
Ji, Yixin
Li, Juntao
Zhang, Min
contents Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning benchmarks. However, their long chain-of-thought reasoning processes incur significant inference overhead. Pruning has emerged as a promising approach to reducing computational costs. However, existing efforts have primarily focused on large language models (LLMs), while pruning LRMs remains unexplored. In this work, we conduct the first empirical study on pruning LRMs and show that directly applying existing pruning techniques fails to yield satisfactory results. Our findings indicate that using self-generated reasoning data for calibration can substantially improve pruning performance. We further investigate how the difficulty and length of reasoning data affect pruning outcomes. Our analysis reveals that challenging and moderately long self-generated reasoning data serve as ideal calibration data. Based on these insights, we propose a Selective Self-Generated Reasoning (SSGR) data construction strategy to provide effective calibration data for pruning LRMs. Experimental results on the DeepSeek-R1-Distill model series validate that our strategy improves the reasoning ability of pruned LRMs by 10%-13% compared to general pruning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
Xiang, Yang
Ji, Yixin
Li, Juntao
Zhang, Min
Computation and Language
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning benchmarks. However, their long chain-of-thought reasoning processes incur significant inference overhead. Pruning has emerged as a promising approach to reducing computational costs. However, existing efforts have primarily focused on large language models (LLMs), while pruning LRMs remains unexplored. In this work, we conduct the first empirical study on pruning LRMs and show that directly applying existing pruning techniques fails to yield satisfactory results. Our findings indicate that using self-generated reasoning data for calibration can substantially improve pruning performance. We further investigate how the difficulty and length of reasoning data affect pruning outcomes. Our analysis reveals that challenging and moderately long self-generated reasoning data serve as ideal calibration data. Based on these insights, we propose a Selective Self-Generated Reasoning (SSGR) data construction strategy to provide effective calibration data for pruning LRMs. Experimental results on the DeepSeek-R1-Distill model series validate that our strategy improves the reasoning ability of pruned LRMs by 10%-13% compared to general pruning methods.
title Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
topic Computation and Language
url https://arxiv.org/abs/2511.18864