Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866911005276110848 |
|---|---|
| author | Liu, Xiaoran He, Siyang Wang, Qiqi Li, Ruixiao Song, Yuerong Liu, Zhigeng Li, Linlin Liu, Qun Huang, Zengfeng Guo, Qipeng He, Ziwei Qiu, Xipeng |
| author_facet | Liu, Xiaoran He, Siyang Wang, Qiqi Li, Ruixiao Song, Yuerong Liu, Zhigeng Li, Linlin Liu, Qun Huang, Zengfeng Guo, Qipeng He, Ziwei Qiu, Xipeng |
| contents | Large Language Models struggle with memory demands from the growing Key-Value (KV) cache as context lengths increase. Existing compression methods homogenize head dimensions or rely on attention-guided token pruning, often sacrificing accuracy or introducing computational overhead. We propose FourierAttention, a training-free framework that exploits the heterogeneous roles of transformer head dimensions: lower dimensions prioritize local context, while upper ones capture long-range dependencies. By projecting the long-context-insensitive dimensions onto orthogonal Fourier bases, FourierAttention approximates their temporal evolution with fixed-length spectral coefficients. Evaluations on LLaMA models show that FourierAttention achieves the best long-context accuracy on LongBench and Needle-In-A-Haystack (NIAH). Besides, a custom Triton kernel, FlashFourierAttention, is designed to optimize memory via streamlined read-write operations, enabling efficient deployment without performance compromise. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_11886 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache Liu, Xiaoran He, Siyang Wang, Qiqi Li, Ruixiao Song, Yuerong Liu, Zhigeng Li, Linlin Liu, Qun Huang, Zengfeng Guo, Qipeng He, Ziwei Qiu, Xipeng Computation and Language Large Language Models struggle with memory demands from the growing Key-Value (KV) cache as context lengths increase. Existing compression methods homogenize head dimensions or rely on attention-guided token pruning, often sacrificing accuracy or introducing computational overhead. We propose FourierAttention, a training-free framework that exploits the heterogeneous roles of transformer head dimensions: lower dimensions prioritize local context, while upper ones capture long-range dependencies. By projecting the long-context-insensitive dimensions onto orthogonal Fourier bases, FourierAttention approximates their temporal evolution with fixed-length spectral coefficients. Evaluations on LLaMA models show that FourierAttention achieves the best long-context accuracy on LongBench and Needle-In-A-Haystack (NIAH). Besides, a custom Triton kernel, FlashFourierAttention, is designed to optimize memory via streamlined read-write operations, enabling efficient deployment without performance compromise. |
| title | Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2506.11886 |