HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Ping, Yi, Jiawei, Wang, Shengnan, Zhang, Juncheng, Jin, Zewen, Zhou, Ouxiang, Liu, Ruibo, Xu, Guanbin, Bai, Youhui, Ye, Bowen, Yuan, Kun, Yang, Tong, Zhang, Gong, Chen, Renhai, Wu, Feng, Li, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
by: Tan, Haoyue, et al.
Published: (2026)
by: Tan, Haoyue, et al.
Published: (2026)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
by: Tzachristas, Georgios, et al.
Published: (2025)
by: Tzachristas, Georgios, et al.
Published: (2025)
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
by: Wang, Shengnan, et al.
Published: (2024)
by: Wang, Shengnan, et al.
Published: (2024)
LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
by: Yi, Jiawei, et al.
Published: (2025)
by: Yi, Jiawei, et al.
Published: (2025)
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
by: Jin, Zewen, et al.
Published: (2025)
by: Jin, Zewen, et al.
Published: (2025)
SpargeAttention2: Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
by: Ai, Xuan, et al.
Published: (2026)
by: Ai, Xuan, et al.
Published: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
by: Xu, Guanbin, et al.
Published: (2026)
by: Xu, Guanbin, et al.
Published: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Mixture-of-Top-k Attention: Efficient Attention via Scalable Fast Weights
by: Wen, Qishuai, et al.
Published: (2026)
by: Wen, Qishuai, et al.
Published: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
TANGNN: a Concise, Scalable and Effective Graph Neural Networks with Top-m Attention Mechanism for Graph Representation Learning
by: E, Jiawei, et al.
Published: (2024)
by: E, Jiawei, et al.
Published: (2024)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
Hardware Implementation of Photonic Spiking Hash Retrieval
by: Shi, Shangxuan, et al.
Published: (2026)
by: Shi, Shangxuan, et al.
Published: (2026)
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Trainable Pointwise Decoder Module for Point Cloud Segmentation
by: Chen, Bike, et al.
Published: (2024)
by: Chen, Bike, et al.
Published: (2024)
TKwinFormer: Top k Window Attention in Vision Transformers for Feature Matching
by: Liao, Yun, et al.
Published: (2023)
by: Liao, Yun, et al.
Published: (2023)
Performance Evaluation of Hashing Algorithms on Commodity Hardware
by: Pandya, Marut
Published: (2024)
by: Pandya, Marut
Published: (2024)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
LATTE: Low-Precision Approximate Attention with Head-wise Trainable Threshold for Efficient Transformer
by: Wang, Jiing-Ping, et al.
Published: (2024)
by: Wang, Jiing-Ping, et al.
Published: (2024)
Gotta Hash 'Em All! Speeding Up Hash Functions for Zero-Knowledge Proof Applications
by: Sheybani, Nojan, et al.
Published: (2025)
by: Sheybani, Nojan, et al.
Published: (2025)
Towards Effective Top-N Hamming Search via Bipartite Graph Contrastive Hashing
by: Chen, Yankai, et al.
Published: (2024)
by: Chen, Yankai, et al.
Published: (2024)
SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
by: Li, Ouxiang, et al.
Published: (2025)
by: Li, Ouxiang, et al.
Published: (2025)
Self-Attention Limits Working Memory Capacity of Transformer-Based Models
by: Gong, Dongyu, et al.
Published: (2024)
by: Gong, Dongyu, et al.
Published: (2024)
ZETA: Leveraging Z-order Curves for Efficient Top-k Attention
by: Zeng, Qiuhao, et al.
Published: (2025)
by: Zeng, Qiuhao, et al.
Published: (2025)
Interpretable descriptors enable prediction of hydrogen-based superconductors at moderate pressures
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Compactness and non-compactness theorems of the fourth- and sixth-order constant $Q$-curvature problems
by: Gong, Liuwei, et al.
Published: (2025)
by: Gong, Liuwei, et al.
Published: (2025)
Global Convergence of the Gursky-Malchiodi $Q$-curvature Flow
by: Gong, Liuwei, et al.
Published: (2026)
by: Gong, Liuwei, et al.
Published: (2026)
UTOPIA: Universally Trainable Optimal Prediction Intervals Aggregation
by: Fan, Jianqing, et al.
Published: (2023)
by: Fan, Jianqing, et al.
Published: (2023)
Scalable Graph Attention-based Instance Selection via Mini-Batch Sampling and Hierarchical Hashing
by: Rustamov, Zahiriddin, et al.
Published: (2025)
by: Rustamov, Zahiriddin, et al.
Published: (2025)
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
by: Sanovar, Rya, et al.
Published: (2024)
by: Sanovar, Rya, et al.
Published: (2024)
Three Generalizations of Erdős Szekeres: $k$-Modal Subsequences
by: Gong, Charles
Published: (2025)
by: Gong, Charles
Published: (2025)
QuaLITi: Quantum Machine Learning Hardware Selection for Inferencing with Top-Tier Performance
by: Phalak, Koustubh, et al.
Published: (2024)
by: Phalak, Koustubh, et al.
Published: (2024)
Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models
by: Qiu, Junxiang, et al.
Published: (2026)
by: Qiu, Junxiang, et al.
Published: (2026)
On the Minimax Regret in Online Ranking with Top-k Feedback
by: Zhang, Mingyuan, et al.
Published: (2023)
by: Zhang, Mingyuan, et al.
Published: (2023)
Engineering Minimal k-Perfect Hash Functions
by: Hermann, Stefan, et al.
Published: (2025)
by: Hermann, Stefan, et al.
Published: (2025)
Similar Items
-
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025) -
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
by: Tan, Haoyue, et al.
Published: (2026) -
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
by: Tzachristas, Georgios, et al.
Published: (2025) -
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
by: Wang, Shengnan, et al.
Published: (2024) -
LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
by: Yi, Jiawei, et al.
Published: (2025)