Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
Fuente:
arXiv
Guardado en:
| Autor principal: | Filipek, Adam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
por: Filipek, Adam
Publicado: (2025)
por: Filipek, Adam
Publicado: (2025)
Coupled Query-Key Dynamics for Attention
por: Gahtan, Barak, et al.
Publicado: (2026)
por: Gahtan, Barak, et al.
Publicado: (2026)
Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
por: Filipek, Adam
Publicado: (2025)
por: Filipek, Adam
Publicado: (2025)
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
por: Le, Dong, et al.
Publicado: (2026)
por: Le, Dong, et al.
Publicado: (2026)
ProxyAttn: Guided Sparse Attention via Representative Heads
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
por: Chen, Yingfa, et al.
Publicado: (2025)
por: Chen, Yingfa, et al.
Publicado: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
por: Musat, Tiberiu
Publicado: (2024)
por: Musat, Tiberiu
Publicado: (2024)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
por: Lai, Xunhao, et al.
Publicado: (2025)
por: Lai, Xunhao, et al.
Publicado: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026)
por: Xu, Ceyu, et al.
Publicado: (2026)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
por: Wang, Hanrui, et al.
Publicado: (2020)
por: Wang, Hanrui, et al.
Publicado: (2020)
SEA: Sparse Linear Attention with Estimated Attention Mask
por: Lee, Heejun, et al.
Publicado: (2023)
por: Lee, Heejun, et al.
Publicado: (2023)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
por: Fan, Zehao, et al.
Publicado: (2025)
por: Fan, Zehao, et al.
Publicado: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
por: Piękos, Piotr, et al.
Publicado: (2025)
por: Piękos, Piotr, et al.
Publicado: (2025)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
por: Souibgui, Mohamed Ali, et al.
Publicado: (2026)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
por: Chen, Feiyang, et al.
Publicado: (2025)
por: Chen, Feiyang, et al.
Publicado: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
por: Chen, Lida, et al.
Publicado: (2025)
por: Chen, Lida, et al.
Publicado: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
por: Hagos, Desta Haileselassie, et al.
Publicado: (2025)
por: Hagos, Desta Haileselassie, et al.
Publicado: (2025)
Block Sparse Flash Attention
por: Ohayon, Daniel, et al.
Publicado: (2025)
por: Ohayon, Daniel, et al.
Publicado: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
por: Lou, Chao, et al.
Publicado: (2024)
por: Lou, Chao, et al.
Publicado: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
por: Jo, Dongwon, et al.
Publicado: (2026)
por: Jo, Dongwon, et al.
Publicado: (2026)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
por: Nawrot, Piotr, et al.
Publicado: (2025)
por: Nawrot, Piotr, et al.
Publicado: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
por: He, Mutian, et al.
Publicado: (2025)
por: He, Mutian, et al.
Publicado: (2025)
Improving Transformers with Dynamically Composable Multi-Head Attention
por: Xiao, Da, et al.
Publicado: (2024)
por: Xiao, Da, et al.
Publicado: (2024)
AdaSplash: Adaptive Sparse Flash Attention
por: Gonçalves, Nuno, et al.
Publicado: (2025)
por: Gonçalves, Nuno, et al.
Publicado: (2025)
Scaling Linear Attention with Sparse State Expansion
por: Pan, Yuqi, et al.
Publicado: (2025)
por: Pan, Yuqi, et al.
Publicado: (2025)
Fast Multipole Attention: A Scalable Multilevel Attention Mechanism for Text and Images
por: Kang, Yanming, et al.
Publicado: (2023)
por: Kang, Yanming, et al.
Publicado: (2023)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
por: Fu, Zizhuo, et al.
Publicado: (2026)
por: Fu, Zizhuo, et al.
Publicado: (2026)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
por: Donhauser, Konstantin, et al.
Publicado: (2025)
por: Donhauser, Konstantin, et al.
Publicado: (2025)
Sparse Attention across Multiple-context KV Cache
por: Cao, Ziyi, et al.
Publicado: (2025)
por: Cao, Ziyi, et al.
Publicado: (2025)
AdaSplash-2: Faster Differentiable Sparse Attention
por: Gonçalves, Nuno, et al.
Publicado: (2026)
por: Gonçalves, Nuno, et al.
Publicado: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
por: Tang, Jiaming, et al.
Publicado: (2024)
por: Tang, Jiaming, et al.
Publicado: (2024)
Generalized Probabilistic Attention Mechanism in Transformers
por: Heo, DongNyeong, et al.
Publicado: (2024)
por: Heo, DongNyeong, et al.
Publicado: (2024)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)
por: Qiu, Quantong, et al.
Publicado: (2026)
Which Attention Heads Matter for In-Context Learning?
por: Yin, Kayo, et al.
Publicado: (2025)
por: Yin, Kayo, et al.
Publicado: (2025)
NOSA: Native and Offloadable Sparse Attention
por: Huang, Yuxiang, et al.
Publicado: (2025)
por: Huang, Yuxiang, et al.
Publicado: (2025)
Ejemplares similares
-
TensorBLEU: Vectorized GPU-based BLEU Score Implementation for Per-Sentence In-Training Evaluation
por: Filipek, Adam
Publicado: (2025) -
Coupled Query-Key Dynamics for Attention
por: Gahtan, Barak, et al.
Publicado: (2026) -
Reactive Transformer (RxT) -- Stateful Real-Time Processing for Event-Driven Reactive Language Models
por: Filipek, Adam
Publicado: (2025) -
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
por: Le, Dong, et al.
Publicado: (2026) -
ProxyAttn: Guided Sparse Attention via Representative Heads
por: Wang, Yixuan, et al.
Publicado: (2025)