SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Anhao, Ye, Fanghua, Fan, Yingqi, Tong, Junlong, Fei, Zhiwei, Su, Hui, Shen, Xiaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913875197165568
author Zhao, Anhao
Ye, Fanghua
Fan, Yingqi
Tong, Junlong
Fei, Zhiwei
Su, Hui
Shen, Xiaoyu
author_facet Zhao, Anhao
Ye, Fanghua
Fan, Yingqi
Tong, Junlong
Fei, Zhiwei
Su, Hui
Shen, Xiaoyu
contents Large language models (LLMs) achieve remarkable performance across tasks but incur substantial computational costs due to their deep, multi-layered architectures. Layer pruning has emerged as a strategy to alleviate these inefficiencies, but conventional static pruning methods overlook two critical dynamics inherent to LLM inference: (1) horizontal dynamics, where token-level heterogeneity demands context-aware pruning decisions, and (2) vertical dynamics, where the distinct functional roles of MLP and self-attention layers necessitate component-specific pruning policies. We introduce SkipGPT, a dynamic layer pruning framework designed to optimize computational resource allocation through two core innovations: (1) global token-aware routing to prioritize critical tokens, and (2) decoupled pruning policies for MLP and self-attention components. To mitigate training instability, we propose a two-stage optimization paradigm: first, a disentangled training phase that learns routing strategies via soft parameterization to avoid premature pruning decisions, followed by parameter-efficient LoRA fine-tuning to restore performance impacted by layer removal. Extensive experiments demonstrate that SkipGPT reduces over 40% of model parameters while matching or exceeding the performance of the original dense model across benchmarks. By harmonizing dynamic efficiency with preserved expressivity, SkipGPT advances the practical deployment of scalable, resource-aware LLMs. Our code is publicly available at: https://github.com/EIT-NLP/SkipGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
Zhao, Anhao
Ye, Fanghua
Fan, Yingqi
Tong, Junlong
Fei, Zhiwei
Su, Hui
Shen, Xiaoyu
Computation and Language
Large language models (LLMs) achieve remarkable performance across tasks but incur substantial computational costs due to their deep, multi-layered architectures. Layer pruning has emerged as a strategy to alleviate these inefficiencies, but conventional static pruning methods overlook two critical dynamics inherent to LLM inference: (1) horizontal dynamics, where token-level heterogeneity demands context-aware pruning decisions, and (2) vertical dynamics, where the distinct functional roles of MLP and self-attention layers necessitate component-specific pruning policies. We introduce SkipGPT, a dynamic layer pruning framework designed to optimize computational resource allocation through two core innovations: (1) global token-aware routing to prioritize critical tokens, and (2) decoupled pruning policies for MLP and self-attention components. To mitigate training instability, we propose a two-stage optimization paradigm: first, a disentangled training phase that learns routing strategies via soft parameterization to avoid premature pruning decisions, followed by parameter-efficient LoRA fine-tuning to restore performance impacted by layer removal. Extensive experiments demonstrate that SkipGPT reduces over 40% of model parameters while matching or exceeding the performance of the original dense model across benchmarks. By harmonizing dynamic efficiency with preserved expressivity, SkipGPT advances the practical deployment of scalable, resource-aware LLMs. Our code is publicly available at: https://github.com/EIT-NLP/SkipGPT.
title SkipGPT: Dynamic Layer Pruning Reinvented with Token Awareness and Module Decoupling
topic Computation and Language
url https://arxiv.org/abs/2506.04179