IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Meng, Wang, Peisong, Shao, Yuantian, Hu, Qinghao, Fang, Hongjian, Zhang, Yifan, Wei, Zhihui, Cheng, Jian
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910008724160512
author Li, Meng
Wang, Peisong
Shao, Yuantian
Hu, Qinghao
Fang, Hongjian
Zhang, Yifan
Wei, Zhihui
Cheng, Jian
author_facet Li, Meng
Wang, Peisong
Shao, Yuantian
Hu, Qinghao
Fang, Hongjian
Zhang, Yifan
Wei, Zhihui
Cheng, Jian
contents Large Language Models (LLMs) achieve strong performance across diverse tasks but face deployment challenges due to their massive size. Structured pruning offers acceleration benefits but leads to significant performance degradation. Recent PCA-based pruning methods have alleviated this issue by retaining key activation components, but are only applied between modules in order to fuse the transformation matrix, which introduces extra parameters and severely disrupts activation distributions due to residual connections. To address these issues, we propose IntraSlice, a framework that applies block-wise module-intra PCA compression pruning. By leveraging the structural characteristics of Transformer modules, we design an approximate PCA method whose transformation matrices can be fully fused into the model without additional parameters. We also introduce a PCA-based global pruning ratio estimator that further considers the distribution of compressed activations, building on conventional module importance. We validate our method on Llama2, Llama3, and Phi series across various language benchmarks. Experimental results demonstrate that our approach achieves superior compression performance compared to recent baselines at the same compression ratio or inference speed.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01975
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
Li, Meng
Wang, Peisong
Shao, Yuantian
Hu, Qinghao
Fang, Hongjian
Zhang, Yifan
Wei, Zhihui
Cheng, Jian
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) achieve strong performance across diverse tasks but face deployment challenges due to their massive size. Structured pruning offers acceleration benefits but leads to significant performance degradation. Recent PCA-based pruning methods have alleviated this issue by retaining key activation components, but are only applied between modules in order to fuse the transformation matrix, which introduces extra parameters and severely disrupts activation distributions due to residual connections. To address these issues, we propose IntraSlice, a framework that applies block-wise module-intra PCA compression pruning. By leveraging the structural characteristics of Transformer modules, we design an approximate PCA method whose transformation matrices can be fully fused into the model without additional parameters. We also introduce a PCA-based global pruning ratio estimator that further considers the distribution of compressed activations, building on conventional module importance. We validate our method on Llama2, Llama3, and Phi series across various language benchmarks. Experimental results demonstrate that our approach achieves superior compression performance compared to recent baselines at the same compression ratio or inference speed.
title IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.01975