FASA: Frequency-aware Sparse Attention

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yifei, Wang, Yueqi, Yue, Zhenrui, Zeng, Huimin, Wang, Yong, Lourentzou, Ismini, Tu, Zhengzhong, Chu, Xiangxiang, McAuley, Julian
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915824117219328
author Wang, Yifei
Wang, Yueqi
Yue, Zhenrui
Zeng, Huimin
Wang, Yong
Lourentzou, Ismini
Tu, Zhengzhong
Chu, Xiangxiang
McAuley, Julian
author_facet Wang, Yifei
Wang, Yueqi
Yue, Zhenrui
Zeng, Huimin
Wang, Yong
Lourentzou, Ismini
Tu, Zhengzhong
Chu, Xiangxiang
McAuley, Julian
contents The deployment of Large Language Models (LLMs) faces a critical bottleneck when handling lengthy inputs: the prohibitive memory footprint of the Key Value (KV) cache. To address this bottleneck, the token pruning paradigm leverages attention sparsity to selectively retain a small, critical subset of tokens. However, existing approaches fall short, with static methods risking irreversible information loss and dynamic strategies employing heuristics that insufficiently capture the query-dependent nature of token importance. We propose FASA, a novel framework that achieves query-aware token eviction by dynamically predicting token importance. FASA stems from a novel insight into RoPE: the discovery of functional sparsity at the frequency-chunk (FC) level. Our key finding is that a small, identifiable subset of "dominant" FCs consistently exhibits high contextual agreement with the full attention head. This provides a robust and computationally free proxy for identifying salient tokens. Building on this insight, FASA first identifies a critical set of tokens using dominant FCs, and then performs focused attention computation solely on this pruned subset. Across a spectrum of long-context tasks, from sequence modeling to complex CoT reasoning, FASA consistently outperforms all token-eviction baselines and achieves near-oracle accuracy, demonstrating remarkable robustness even under constraint budgets. Notably, on LongBench-V1, FASA reaches nearly 100\% of full-KV performance when only keeping 256 tokens, and achieves 2.56$\times$ speedup using just 18.9\% of the cache on AIME24.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03152
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FASA: Frequency-aware Sparse Attention
Wang, Yifei
Wang, Yueqi
Yue, Zhenrui
Zeng, Huimin
Wang, Yong
Lourentzou, Ismini
Tu, Zhengzhong
Chu, Xiangxiang
McAuley, Julian
Computation and Language
The deployment of Large Language Models (LLMs) faces a critical bottleneck when handling lengthy inputs: the prohibitive memory footprint of the Key Value (KV) cache. To address this bottleneck, the token pruning paradigm leverages attention sparsity to selectively retain a small, critical subset of tokens. However, existing approaches fall short, with static methods risking irreversible information loss and dynamic strategies employing heuristics that insufficiently capture the query-dependent nature of token importance. We propose FASA, a novel framework that achieves query-aware token eviction by dynamically predicting token importance. FASA stems from a novel insight into RoPE: the discovery of functional sparsity at the frequency-chunk (FC) level. Our key finding is that a small, identifiable subset of "dominant" FCs consistently exhibits high contextual agreement with the full attention head. This provides a robust and computationally free proxy for identifying salient tokens. Building on this insight, FASA first identifies a critical set of tokens using dominant FCs, and then performs focused attention computation solely on this pruned subset. Across a spectrum of long-context tasks, from sequence modeling to complex CoT reasoning, FASA consistently outperforms all token-eviction baselines and achieves near-oracle accuracy, demonstrating remarkable robustness even under constraint budgets. Notably, on LongBench-V1, FASA reaches nearly 100\% of full-KV performance when only keeping 256 tokens, and achieves 2.56$\times$ speedup using just 18.9\% of the cache on AIME24.
title FASA: Frequency-aware Sparse Attention
topic Computation and Language
url https://arxiv.org/abs/2602.03152