DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Ruokai, Li, Yuhang, Lee, Donghyun, Panda, Priyadarshini
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912706116714496
author Yin, Ruokai
Li, Yuhang
Lee, Donghyun
Panda, Priyadarshini
author_facet Yin, Ruokai
Li, Yuhang
Lee, Donghyun
Panda, Priyadarshini
contents Large language models (LLMs) deliver strong performance but are difficult to deploy due to high memory and compute costs. While pruning reduces these demands, most methods ignore activation sparsity observed at runtime. We reinterpret activation sparsity as dynamic structured weight sparsity and propose DuoGPT, a unified framework that constructs dual-sparse (spMspV) workloads by combining unstructured weight pruning with activation sparsity. To preserve accuracy, we extend the Optimal Brain Compression (OBC) framework with activation-aware calibration and introduce output residuals from the dense model as correction terms. We further optimize the solution for efficient GPU execution, enabling scalability to billion-parameter LLMs. Evaluations on LLaMA-2 and LLaMA-3 show that DuoGPT outperforms state-of-the-art structured pruning methods by up to 9.17% accuracy at an iso-speedup of 1.39$\times$ compared to the baseline dense model. Code is available at Github.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20194
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
Yin, Ruokai
Li, Yuhang
Lee, Donghyun
Panda, Priyadarshini
Machine Learning
Large language models (LLMs) deliver strong performance but are difficult to deploy due to high memory and compute costs. While pruning reduces these demands, most methods ignore activation sparsity observed at runtime. We reinterpret activation sparsity as dynamic structured weight sparsity and propose DuoGPT, a unified framework that constructs dual-sparse (spMspV) workloads by combining unstructured weight pruning with activation sparsity. To preserve accuracy, we extend the Optimal Brain Compression (OBC) framework with activation-aware calibration and introduce output residuals from the dense model as correction terms. We further optimize the solution for efficient GPU execution, enabling scalability to billion-parameter LLMs. Evaluations on LLaMA-2 and LLaMA-3 show that DuoGPT outperforms state-of-the-art structured pruning methods by up to 9.17% accuracy at an iso-speedup of 1.39$\times$ compared to the baseline dense model. Code is available at Github.
title DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
topic Machine Learning
url https://arxiv.org/abs/2506.20194