Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oh, Jungsuk, Jeon, Hyeseo, Ji, Hyunjune, Kong, Kyongmin, Lee, Jay-Yoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
KV Admission: Learning What to Write for Efficient Long-Context Inference
von: Huang, Yen-Chieh, et al.
Veröffentlicht: (2025)
von: Huang, Yen-Chieh, et al.
Veröffentlicht: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
von: Li, Kunxi, et al.
Veröffentlicht: (2025)
Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
von: Peng, Dan, et al.
Veröffentlicht: (2025)
von: Peng, Dan, et al.
Veröffentlicht: (2025)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
von: Li, Junjie, et al.
Veröffentlicht: (2026)
von: Li, Junjie, et al.
Veröffentlicht: (2026)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
Training-Free Exponential Context Extension via Cascading KV Cache
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
von: Willette, Jeffrey, et al.
Veröffentlicht: (2024)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
von: Li, Dongfang, et al.
Veröffentlicht: (2026)
MatKV: Trading Compute for Flash Storage in LLM Inference
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
von: Shin, Kun-Woo, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
von: Guanzhong, Chen
Veröffentlicht: (2026)
von: Guanzhong, Chen
Veröffentlicht: (2026)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
von: Gao, Yifei, et al.
Veröffentlicht: (2026)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
POP: Prefill-Only Pruning for Efficient Large Model Inference
von: He, Junhui, et al.
Veröffentlicht: (2026)
von: He, Junhui, et al.
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
von: Tao, Qian, et al.
Veröffentlicht: (2024)
von: Tao, Qian, et al.
Veröffentlicht: (2024)
Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge Distillation
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
von: Shin, Hyunjune, et al.
Veröffentlicht: (2024)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
von: Oh, Jungsuk, et al.
Veröffentlicht: (2025) -
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
von: Jones, Dalton, et al.
Veröffentlicht: (2026) -
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026) -
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
von: Tang, Xiaojuan, et al.
Veröffentlicht: (2025) -
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
von: Zhang, Hao, et al.
Veröffentlicht: (2025)