Towards Universal & Efficient Model Compression via Exponential Torque Pruning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Modi, Sarthak Ketanbhai, Lim, Zi Pong, Kuchhal, Shourya, Cao, Yushi, Cheng, Yupeng, Teo, Yon Shin, Lin, Shang-Wei, Li, Zhiming
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909674000875520
author Modi, Sarthak Ketanbhai
Lim, Zi Pong
Kuchhal, Shourya
Cao, Yushi
Cheng, Yupeng
Teo, Yon Shin
Lin, Shang-Wei
Li, Zhiming
author_facet Modi, Sarthak Ketanbhai
Lim, Zi Pong
Kuchhal, Shourya
Cao, Yushi
Cheng, Yupeng
Teo, Yon Shin
Lin, Shang-Wei
Li, Zhiming
contents The rapid growth in complexity and size of modern deep neural networks (DNNs) has increased challenges related to computational costs and memory usage, spurring a growing interest in efficient model compression techniques. Previous state-of-the-art approach proposes using a Torque-inspired regularization which forces the weights of neural modules around a selected pivot point. Whereas, we observe that the pruning effect of this approach is far from perfect, as the post-trained network is still dense and also suffers from high accuracy drop. In this work, we attribute such ineffectiveness to the default linear force application scheme, which imposes inappropriate force on neural module of different distances. To efficiently prune the redundant and distant modules while retaining those that are close and necessary for effective inference, in this work, we propose Exponential Torque Pruning (ETP), which adopts an exponential force application scheme for regularization. Experimental results on a broad range of domains demonstrate that, though being extremely simple, ETP manages to achieve significantly higher compression rate than the previous state-of-the-art pruning strategies with negligible accuracy drop.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22015
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Universal & Efficient Model Compression via Exponential Torque Pruning
Modi, Sarthak Ketanbhai
Lim, Zi Pong
Kuchhal, Shourya
Cao, Yushi
Cheng, Yupeng
Teo, Yon Shin
Lin, Shang-Wei
Li, Zhiming
Computer Vision and Pattern Recognition
The rapid growth in complexity and size of modern deep neural networks (DNNs) has increased challenges related to computational costs and memory usage, spurring a growing interest in efficient model compression techniques. Previous state-of-the-art approach proposes using a Torque-inspired regularization which forces the weights of neural modules around a selected pivot point. Whereas, we observe that the pruning effect of this approach is far from perfect, as the post-trained network is still dense and also suffers from high accuracy drop. In this work, we attribute such ineffectiveness to the default linear force application scheme, which imposes inappropriate force on neural module of different distances. To efficiently prune the redundant and distant modules while retaining those that are close and necessary for effective inference, in this work, we propose Exponential Torque Pruning (ETP), which adopts an exponential force application scheme for regularization. Experimental results on a broad range of domains demonstrate that, though being extremely simple, ETP manages to achieve significantly higher compression rate than the previous state-of-the-art pruning strategies with negligible accuracy drop.
title Towards Universal & Efficient Model Compression via Exponential Torque Pruning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.22015