HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Minghui, Rabbani, Tahseen, O'Halloran, Tony, Sankaralingam, Ananth, Hartley, Mary-Anne, Huang, Furong, Fermüller, Cornelia, Aloimonos, Yiannis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023)
by: Yuan, Dehao, et al.
Published: (2023)
PHast -- Perfect Hashing made fast
by: Beling, Piotr, et al.
Published: (2025)
by: Beling, Piotr, et al.
Published: (2025)
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024)
by: Devaraj, Chinmaya, et al.
Published: (2024)
Maelstrom Networks
by: Evanusa, Matthew, et al.
Published: (2024)
by: Evanusa, Matthew, et al.
Published: (2024)
Embodied Visuomotor Representation
by: Burner, Levi, et al.
Published: (2024)
by: Burner, Levi, et al.
Published: (2024)
Local Rendezvous Hashing: Bounded Loads and Minimal Churn via Cache-Local Candidates
by: Guan, Yongjie
Published: (2025)
by: Guan, Yongjie
Published: (2025)
Evaluation of Hash Algorithm Performance for Cryptocurrency Exchanges Based on Blockchain System
by: Chen, Abel C. H.
Published: (2024)
by: Chen, Abel C. H.
Published: (2024)
eHashPipe: Lightweight Top-K and Per-PID Resource Monitoring with eBPF
by: Dai, Yuanjun, et al.
Published: (2025)
by: Dai, Yuanjun, et al.
Published: (2025)
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
by: Wu, Jiayi, et al.
Published: (2026)
by: Wu, Jiayi, et al.
Published: (2026)
MapReplay: Trace-Driven Benchmark Generation for Java HashMap
by: Schiavio, Filippo, et al.
Published: (2026)
by: Schiavio, Filippo, et al.
Published: (2026)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
by: Zhang, Quqing, et al.
Published: (2026)
by: Zhang, Quqing, et al.
Published: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
by: Liu, Guangda, et al.
Published: (2024)
by: Liu, Guangda, et al.
Published: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
by: Liu, Hongyao, et al.
Published: (2026)
by: Liu, Hongyao, et al.
Published: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
by: Taneja, Maanas, et al.
Published: (2026)
by: Taneja, Maanas, et al.
Published: (2026)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-scale Complex Unknown Environments
by: He, Botao, et al.
Published: (2024)
by: He, Botao, et al.
Published: (2024)
ViewActive: Active viewpoint optimization from a single image
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
by: Barad, Haim, et al.
Published: (2023)
by: Barad, Haim, et al.
Published: (2023)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
Learning Normal Flow Directly From Event Neighborhoods
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
by: Jiang, Chaoyi, et al.
Published: (2024)
by: Jiang, Chaoyi, et al.
Published: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
Active Human Pose Estimation via an Autonomous UAV Agent
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
CAMP: A Cost Adaptive Multi-Queue Eviction Policy for Key-Value Stores
by: Ghandeharizadeh, Shahram, et al.
Published: (2024)
by: Ghandeharizadeh, Shahram, et al.
Published: (2024)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
by: Guo, Jinyu, et al.
Published: (2026)
by: Guo, Jinyu, et al.
Published: (2026)
In-context KV-Cache Eviction for LLMs via Attention-Gate
by: Zeng, Zihao, et al.
Published: (2024)
by: Zeng, Zihao, et al.
Published: (2024)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)
by: Yang, Chongzhuo, et al.
Published: (2024)
Similar Items
-
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
by: Yuan, Dehao, et al.
Published: (2024) -
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025) -
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023) -
PHast -- Perfect Hashing made fast
by: Beling, Piotr, et al.
Published: (2025) -
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024)