Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Xiaoran, He, Siyang, Wang, Qiqi, Li, Ruixiao, Song, Yuerong, Liu, Zhigeng, Li, Linlin, Liu, Qun, Huang, Zengfeng, Guo, Qipeng, He, Ziwei, Qiu, Xipeng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911005276110848
author Liu, Xiaoran
He, Siyang
Wang, Qiqi
Li, Ruixiao
Song, Yuerong
Liu, Zhigeng
Li, Linlin
Liu, Qun
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
author_facet Liu, Xiaoran
He, Siyang
Wang, Qiqi
Li, Ruixiao
Song, Yuerong
Liu, Zhigeng
Li, Linlin
Liu, Qun
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
contents Large Language Models struggle with memory demands from the growing Key-Value (KV) cache as context lengths increase. Existing compression methods homogenize head dimensions or rely on attention-guided token pruning, often sacrificing accuracy or introducing computational overhead. We propose FourierAttention, a training-free framework that exploits the heterogeneous roles of transformer head dimensions: lower dimensions prioritize local context, while upper ones capture long-range dependencies. By projecting the long-context-insensitive dimensions onto orthogonal Fourier bases, FourierAttention approximates their temporal evolution with fixed-length spectral coefficients. Evaluations on LLaMA models show that FourierAttention achieves the best long-context accuracy on LongBench and Needle-In-A-Haystack (NIAH). Besides, a custom Triton kernel, FlashFourierAttention, is designed to optimize memory via streamlined read-write operations, enabling efficient deployment without performance compromise.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
Liu, Xiaoran
He, Siyang
Wang, Qiqi
Li, Ruixiao
Song, Yuerong
Liu, Zhigeng
Li, Linlin
Liu, Qun
Huang, Zengfeng
Guo, Qipeng
He, Ziwei
Qiu, Xipeng
Computation and Language
Large Language Models struggle with memory demands from the growing Key-Value (KV) cache as context lengths increase. Existing compression methods homogenize head dimensions or rely on attention-guided token pruning, often sacrificing accuracy or introducing computational overhead. We propose FourierAttention, a training-free framework that exploits the heterogeneous roles of transformer head dimensions: lower dimensions prioritize local context, while upper ones capture long-range dependencies. By projecting the long-context-insensitive dimensions onto orthogonal Fourier bases, FourierAttention approximates their temporal evolution with fixed-length spectral coefficients. Evaluations on LLaMA models show that FourierAttention achieves the best long-context accuracy on LongBench and Needle-In-A-Haystack (NIAH). Besides, a custom Triton kernel, FlashFourierAttention, is designed to optimize memory via streamlined read-write operations, enabling efficient deployment without performance compromise.
title Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
topic Computation and Language
url https://arxiv.org/abs/2506.11886