Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuxiang, Yuan, Binhang, Han, Xu, Xiao, Chaojun, Liu, Zhiyuan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026)
by: Yang, Xintong, et al.
Published: (2026)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Taming the Fragility of KV Cache Eviction in LLM Inference
by: Feng, Yuan, et al.
Published: (2025)
by: Feng, Yuan, et al.
Published: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Mask Tokens as Prophet: Fine-Grained Cache Eviction for Efficient dLLM Inference
by: Huang, Jianuo, et al.
Published: (2025)
by: Huang, Jianuo, et al.
Published: (2025)
Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer Inference
by: Fei, Weizhi, et al.
Published: (2025)
by: Fei, Weizhi, et al.
Published: (2025)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
by: Kim, Jang-Hyun, et al.
Published: (2026)
by: Kim, Jang-Hyun, et al.
Published: (2026)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial Contexts
by: Ge, Suyu, et al.
Published: (2024)
by: Ge, Suyu, et al.
Published: (2024)
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
by: Feng, Yuan, et al.
Published: (2024)
by: Feng, Yuan, et al.
Published: (2024)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
by: Zhou, Enshuai, et al.
Published: (2026)
by: Zhou, Enshuai, et al.
Published: (2026)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
by: Song, Yuerong, et al.
Published: (2025)
by: Song, Yuerong, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
by: Lin, Gang, et al.
Published: (2026)
by: Lin, Gang, et al.
Published: (2026)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
by: Svirschevski, Ruslan, et al.
Published: (2024)
by: Svirschevski, Ruslan, et al.
Published: (2024)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
Retrieval Head Mechanistically Explains Long-Context Factuality
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
by: Bai, Yu, et al.
Published: (2024)
by: Bai, Yu, et al.
Published: (2024)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
by: Liu, Runheng, et al.
Published: (2024)
by: Liu, Runheng, et al.
Published: (2024)
Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilities
by: Padmanabhan, Shankar, et al.
Published: (2026)
by: Padmanabhan, Shankar, et al.
Published: (2026)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
by: Long, Lingkun, et al.
Published: (2025)
by: Long, Lingkun, et al.
Published: (2025)
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
by: Yao, Jinwei, et al.
Published: (2024)
by: Yao, Jinwei, et al.
Published: (2024)
CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices
by: Zhao, Weilin, et al.
Published: (2023)
by: Zhao, Weilin, et al.
Published: (2023)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Similar Items
-
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024) -
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025) -
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026) -
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026) -
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)