SAP: Syntactic Attention Pruning for Transformer-based Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Tzu-Yun, Hong, Ding-Yong, Wu, Jan-Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912782180417536
author Lee, Tzu-Yun
Hong, Ding-Yong
Wu, Jan-Jan
author_facet Lee, Tzu-Yun
Hong, Ding-Yong
Wu, Jan-Jan
contents This paper introduces Syntactic Attention Pruning (SAP), a novel method for effectively pruning attention heads in Transformer models. Unlike conventional approaches that rely solely on mathematical analysis of model weights and activations, SAP incorporates both the syntactic structure and attention patterns of sentences to guide the pruning process. By leveraging these linguistic features, SAP not only achieves performance comparable to state-of-the-art methods but also enhances the interpretability of model behavior. To further improve robustness, we propose Candidate Filtering (CF), a mechanism that prioritizes heads based on their contribution to model performance, mitigating degradation during pruning. Experimental results indicate that SAP effectively preserves critical heads of a high density of strong attention values, outperforming existing head pruning strategies in retrain-free settings. These findings position SAP as a promising foundation for a new direction in model compression research, offering high flexibility for pruning across all transformer-based language models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAP: Syntactic Attention Pruning for Transformer-based Language Models
Lee, Tzu-Yun
Hong, Ding-Yong
Wu, Jan-Jan
Computation and Language
Machine Learning
This paper introduces Syntactic Attention Pruning (SAP), a novel method for effectively pruning attention heads in Transformer models. Unlike conventional approaches that rely solely on mathematical analysis of model weights and activations, SAP incorporates both the syntactic structure and attention patterns of sentences to guide the pruning process. By leveraging these linguistic features, SAP not only achieves performance comparable to state-of-the-art methods but also enhances the interpretability of model behavior. To further improve robustness, we propose Candidate Filtering (CF), a mechanism that prioritizes heads based on their contribution to model performance, mitigating degradation during pruning. Experimental results indicate that SAP effectively preserves critical heads of a high density of strong attention values, outperforming existing head pruning strategies in retrain-free settings. These findings position SAP as a promising foundation for a new direction in model compression research, offering high flexibility for pruning across all transformer-based language models.
title SAP: Syntactic Attention Pruning for Transformer-based Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.19125