Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alim, Md. Samiul, Khan, Sharjil, Biswas, Amrijit, Rahman, Fuad, Rahman, Shafin, Mohammed, Nabeel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908667633205248
author Alim, Md. Samiul
Khan, Sharjil
Biswas, Amrijit
Rahman, Fuad
Rahman, Shafin
Mohammed, Nabeel
author_facet Alim, Md. Samiul
Khan, Sharjil
Biswas, Amrijit
Rahman, Fuad
Rahman, Shafin
Mohammed, Nabeel
contents Unstructured pruning remains a powerful strategy for compressing deep neural networks, yet it often demands iterative train-prune-retrain cycles, resulting in significant computational overhead. To address this challenge, we introduce a novel teacher-guided pruning framework that tightly integrates Knowledge Distillation (KD) with importance score estimation. Unlike prior approaches that apply KD as a post-pruning recovery step, our method leverages gradient signals informed by the teacher during importance score calculation to identify and retain parameters most critical for both task performance and knowledge transfer. Our method facilitates a one-shot global pruning strategy that efficiently eliminates redundant weights while preserving essential representations. After pruning, we employ sparsity-aware retraining with and without KD to recover accuracy without reactivating pruned connections. Comprehensive experiments across multiple image classification benchmarks, including CIFAR-10, CIFAR-100, and TinyImageNet, demonstrate that our method consistently achieves high sparsity levels with minimal performance degradation. Notably, our approach outperforms state-of-the-art baselines such as EPG and EPSD at high sparsity levels, while offering a more computationally efficient alternative to iterative pruning schemes like COLT. The proposed framework offers a computation-efficient, performance-preserving solution well suited for deployment in resource-constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16653
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
Alim, Md. Samiul
Khan, Sharjil
Biswas, Amrijit
Rahman, Fuad
Rahman, Shafin
Mohammed, Nabeel
Computer Vision and Pattern Recognition
Artificial Intelligence
Unstructured pruning remains a powerful strategy for compressing deep neural networks, yet it often demands iterative train-prune-retrain cycles, resulting in significant computational overhead. To address this challenge, we introduce a novel teacher-guided pruning framework that tightly integrates Knowledge Distillation (KD) with importance score estimation. Unlike prior approaches that apply KD as a post-pruning recovery step, our method leverages gradient signals informed by the teacher during importance score calculation to identify and retain parameters most critical for both task performance and knowledge transfer. Our method facilitates a one-shot global pruning strategy that efficiently eliminates redundant weights while preserving essential representations. After pruning, we employ sparsity-aware retraining with and without KD to recover accuracy without reactivating pruned connections. Comprehensive experiments across multiple image classification benchmarks, including CIFAR-10, CIFAR-100, and TinyImageNet, demonstrate that our method consistently achieves high sparsity levels with minimal performance degradation. Notably, our approach outperforms state-of-the-art baselines such as EPG and EPSD at high sparsity levels, while offering a more computationally efficient alternative to iterative pruning schemes like COLT. The proposed framework offers a computation-efficient, performance-preserving solution well suited for deployment in resource-constrained environments.
title Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.16653