Interleaved Head Attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Duvvuri, Sai Surya, Ekbote, Chanakya, Bansal, Rachit, Tiwari, Rishabh, Khatri, Devvrit, Brandfonbrener, David, Liang, Paul, Dhillon, Inderjit, Zaheer, Manzil
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912925721034752
author Duvvuri, Sai Surya
Ekbote, Chanakya
Bansal, Rachit
Tiwari, Rishabh
Khatri, Devvrit
Brandfonbrener, David
Liang, Paul
Dhillon, Inderjit
Zaheer, Manzil
author_facet Duvvuri, Sai Surya
Ekbote, Chanakya
Bansal, Rachit
Tiwari, Rishabh
Khatri, Devvrit
Brandfonbrener, David
Liang, Paul
Dhillon, Inderjit
Zaheer, Manzil
contents Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, MHA suffers from a fundamental linear scaling limitation: $H$ attention heads produce exactly $H$ independent attention matrices, with no communication between heads during attention computation. This becomes problematic for multi-step reasoning, where correct answers depend on aggregating evidence from multiple parts of the context and composing latent token-to-token relations over a chain of intermediate inferences. To address this, we propose Interleaved Head Attention (IHA), which enables cross-head mixing by constructing $P$ pseudo-heads per head (typically $P=H$), where each pseudo query/key/value is a learned linear combination of all $H$ original queries, keys and values respectively. Interactions between pseudo-query and pseudo-key heads induce up to $P^2$ attention patterns per head with modest parameter overhead $\mathcal{O}(H^2P)$. We provide theory showing improved efficiency in terms of number of parameters on the synthetic Polynomial task (IHA uses $Θ(\sqrt{k}n^2)$ parameters vs. $Θ(kn^2)$ for MHA) and on the synthetic order-sensitive CPM-3 task (IHA uses $\lceil\sqrt{N_{\max}}\rceil$ heads vs. $N_{\max}$ for MHA). On real-world benchmarks, IHA improves Multi-Key retrieval on RULER by 10-20% (4k-16k) and, after fine-tuning for reasoning on OpenThoughts, improves GSM8K by 5.8% and MATH-500 by 2.8% (Majority Vote) over full attention.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21371
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Interleaved Head Attention
Duvvuri, Sai Surya
Ekbote, Chanakya
Bansal, Rachit
Tiwari, Rishabh
Khatri, Devvrit
Brandfonbrener, David
Liang, Paul
Dhillon, Inderjit
Zaheer, Manzil
Machine Learning
Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, MHA suffers from a fundamental linear scaling limitation: $H$ attention heads produce exactly $H$ independent attention matrices, with no communication between heads during attention computation. This becomes problematic for multi-step reasoning, where correct answers depend on aggregating evidence from multiple parts of the context and composing latent token-to-token relations over a chain of intermediate inferences. To address this, we propose Interleaved Head Attention (IHA), which enables cross-head mixing by constructing $P$ pseudo-heads per head (typically $P=H$), where each pseudo query/key/value is a learned linear combination of all $H$ original queries, keys and values respectively. Interactions between pseudo-query and pseudo-key heads induce up to $P^2$ attention patterns per head with modest parameter overhead $\mathcal{O}(H^2P)$. We provide theory showing improved efficiency in terms of number of parameters on the synthetic Polynomial task (IHA uses $Θ(\sqrt{k}n^2)$ parameters vs. $Θ(kn^2)$ for MHA) and on the synthetic order-sensitive CPM-3 task (IHA uses $\lceil\sqrt{N_{\max}}\rceil$ heads vs. $N_{\max}$ for MHA). On real-world benchmarks, IHA improves Multi-Key retrieval on RULER by 10-20% (4k-16k) and, after fine-tuning for reasoning on OpenThoughts, improves GSM8K by 5.8% and MATH-500 by 2.8% (Majority Vote) over full attention.
title Interleaved Head Attention
topic Machine Learning
url https://arxiv.org/abs/2602.21371