Hopscotch: Discovering and Skipping Redundancies in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912586957586432 |
|---|---|
| author | Eyceoz, Mustafa Nayak, Nikhil Shivakumar Wang, Hao Han, Ligong Srivastava, Akash |
| author_facet | Eyceoz, Mustafa Nayak, Nikhil Shivakumar Wang, Hao Han, Ligong Srivastava, Akash |
| contents | Modern causal language models stack many attention blocks to improve performance, but not all blocks are necessary for every task. We propose Hopscotch, a simple yet effective method that identifies and skips attention blocks with least contributions to a task and adapts to preserve output quality. Hopscotch jointly optimizes which blocks to skip and how to scale the outputs of the remaining layers. By introducing lightweight, trainable scaling parameters to attention and MLP blocks, it mitigates distribution shifts in hidden states caused by removing attention blocks. Hopscotch does not modify model weights or require access to pretraining or instruction-tuning data, and is compatible with existing model compression techniques. When applied to $\texttt{Llama-3.1-8B}$ and $\texttt{Qwen2.5-7B}$, Hopscotch achieves less than a 2% drop in performance even after skipping four attention blocks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_03303 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Hopscotch: Discovering and Skipping Redundancies in Language Models Eyceoz, Mustafa Nayak, Nikhil Shivakumar Wang, Hao Han, Ligong Srivastava, Akash Computation and Language Artificial Intelligence Machine Learning 68T50 I.2.7; I.2.6; I.2.4 Modern causal language models stack many attention blocks to improve performance, but not all blocks are necessary for every task. We propose Hopscotch, a simple yet effective method that identifies and skips attention blocks with least contributions to a task and adapts to preserve output quality. Hopscotch jointly optimizes which blocks to skip and how to scale the outputs of the remaining layers. By introducing lightweight, trainable scaling parameters to attention and MLP blocks, it mitigates distribution shifts in hidden states caused by removing attention blocks. Hopscotch does not modify model weights or require access to pretraining or instruction-tuning data, and is compatible with existing model compression techniques. When applied to $\texttt{Llama-3.1-8B}$ and $\texttt{Qwen2.5-7B}$, Hopscotch achieves less than a 2% drop in performance even after skipping four attention blocks. |
| title | Hopscotch: Discovering and Skipping Redundancies in Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning 68T50 I.2.7; I.2.6; I.2.4 |
| url | https://arxiv.org/abs/2506.03303 |