Hopscotch: Discovering and Skipping Redundancies in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Eyceoz, Mustafa, Nayak, Nikhil Shivakumar, Wang, Hao, Han, Ligong, Srivastava, Akash
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912586957586432
author Eyceoz, Mustafa
Nayak, Nikhil Shivakumar
Wang, Hao
Han, Ligong
Srivastava, Akash
author_facet Eyceoz, Mustafa
Nayak, Nikhil Shivakumar
Wang, Hao
Han, Ligong
Srivastava, Akash
contents Modern causal language models stack many attention blocks to improve performance, but not all blocks are necessary for every task. We propose Hopscotch, a simple yet effective method that identifies and skips attention blocks with least contributions to a task and adapts to preserve output quality. Hopscotch jointly optimizes which blocks to skip and how to scale the outputs of the remaining layers. By introducing lightweight, trainable scaling parameters to attention and MLP blocks, it mitigates distribution shifts in hidden states caused by removing attention blocks. Hopscotch does not modify model weights or require access to pretraining or instruction-tuning data, and is compatible with existing model compression techniques. When applied to $\texttt{Llama-3.1-8B}$ and $\texttt{Qwen2.5-7B}$, Hopscotch achieves less than a 2% drop in performance even after skipping four attention blocks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hopscotch: Discovering and Skipping Redundancies in Language Models
Eyceoz, Mustafa
Nayak, Nikhil Shivakumar
Wang, Hao
Han, Ligong
Srivastava, Akash
Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7; I.2.6; I.2.4
Modern causal language models stack many attention blocks to improve performance, but not all blocks are necessary for every task. We propose Hopscotch, a simple yet effective method that identifies and skips attention blocks with least contributions to a task and adapts to preserve output quality. Hopscotch jointly optimizes which blocks to skip and how to scale the outputs of the remaining layers. By introducing lightweight, trainable scaling parameters to attention and MLP blocks, it mitigates distribution shifts in hidden states caused by removing attention blocks. Hopscotch does not modify model weights or require access to pretraining or instruction-tuning data, and is compatible with existing model compression techniques. When applied to $\texttt{Llama-3.1-8B}$ and $\texttt{Qwen2.5-7B}$, Hopscotch achieves less than a 2% drop in performance even after skipping four attention blocks.
title Hopscotch: Discovering and Skipping Redundancies in Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
68T50
I.2.7; I.2.6; I.2.4
url https://arxiv.org/abs/2506.03303