QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jones, Dalton, Park, Junyoung, Morse, Matthew, Lee, Mingu, Lott, Chris, Langston, Harper |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
von: Gautam, Aayush, et al.
Veröffentlicht: (2026)
von: Gautam, Aayush, et al.
Veröffentlicht: (2026)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
KVComm: Enabling Efficient LLM Communication through Selective KV Sharing
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Shi, Xiangyu, et al.
Veröffentlicht: (2025)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
von: Wang, Meixuan, et al.
Veröffentlicht: (2025)
von: Wang, Meixuan, et al.
Veröffentlicht: (2025)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuxuan, et al.
Veröffentlicht: (2025)
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data
von: Kang, Mingu, et al.
Veröffentlicht: (2024)
von: Kang, Mingu, et al.
Veröffentlicht: (2024)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
von: An, Tai, et al.
Veröffentlicht: (2025)
von: An, Tai, et al.
Veröffentlicht: (2025)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
von: Li, Junjie, et al.
Veröffentlicht: (2026)
von: Li, Junjie, et al.
Veröffentlicht: (2026)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
IntroLM: Introspective Language Models via Prefilling-Time Self-Evaluation
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
von: Woo, Sunghyeon, et al.
Veröffentlicht: (2026)
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2026)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
von: Qiu, Ruizhong, et al.
Veröffentlicht: (2024)
MatKV: Trading Compute for Flash Storage in LLM Inference
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
CASK: Core-Aware Selective KV Compression for Reasoning Traces
von: Kim, Buseong, et al.
Veröffentlicht: (2026)
von: Kim, Buseong, et al.
Veröffentlicht: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
Joint Encoding of KV-Cache Blocks for Scalable LLM Serving
von: Kampeas, Joseph, et al.
Veröffentlicht: (2026)
von: Kampeas, Joseph, et al.
Veröffentlicht: (2026)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
von: Guanzhong, Chen
Veröffentlicht: (2026)
von: Guanzhong, Chen
Veröffentlicht: (2026)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
von: Kang, Hao, et al.
Veröffentlicht: (2024)
von: Kang, Hao, et al.
Veröffentlicht: (2024)
Shaping Zero-Shot Coordination via State Blocking
von: Kang, Mingu, et al.
Veröffentlicht: (2026)
von: Kang, Mingu, et al.
Veröffentlicht: (2026)
Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning
von: Lee, Sunwoo, et al.
Veröffentlicht: (2026)
von: Lee, Sunwoo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025) -
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
von: Gautam, Aayush, et al.
Veröffentlicht: (2026) -
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
von: Park, Junyoung, et al.
Veröffentlicht: (2025) -
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024) -
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)