Saved in:
| Main Authors: | Gim, In, Chen, Guojun, Lee, Seung-seob, Sarda, Nikhil, Khandelwal, Anurag, Zhong, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.04934 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asynchronous LLM Function Calling
by: Gim, In, et al.
Published: (2024)
by: Gim, In, et al.
Published: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
by: Ma, Zhiyao, et al.
Published: (2025)
by: Ma, Zhiyao, et al.
Published: (2025)
Pie: A Programmable Serving System for Emerging LLM Applications
by: Gim, In, et al.
Published: (2025)
by: Gim, In, et al.
Published: (2025)
LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
by: Bansal, Harsh Vardhan
Published: (2025)
by: Bansal, Harsh Vardhan
Published: (2025)
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
by: Su, Jinbo, et al.
Published: (2026)
by: Su, Jinbo, et al.
Published: (2026)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
by: Li, Weiqing, et al.
Published: (2025)
by: Li, Weiqing, et al.
Published: (2025)
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
by: Lee, Kun-Hui, et al.
Published: (2025)
by: Lee, Kun-Hui, et al.
Published: (2025)
PULSE: Accelerating Distributed Pointer-Traversals on Disaggregated Memory (Extended Version)
by: Tang, Yupeng, et al.
Published: (2023)
by: Tang, Yupeng, et al.
Published: (2023)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse
by: Yang, Huan, et al.
Published: (2025)
by: Yang, Huan, et al.
Published: (2025)
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
by: Kummer, Cornelius, et al.
Published: (2026)
by: Kummer, Cornelius, et al.
Published: (2026)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
by: Monteiro, João, et al.
Published: (2024)
by: Monteiro, João, et al.
Published: (2024)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
by: Agarwal, Utkarsh, et al.
Published: (2024)
by: Agarwal, Utkarsh, et al.
Published: (2024)
Generative Caching for Structurally Similar Prompts and Responses
by: Chakraborty, Sarthak, et al.
Published: (2025)
by: Chakraborty, Sarthak, et al.
Published: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025)
by: Gu, Yifeng, et al.
Published: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
Serve Programs, Not Prompts
by: Gim, In, et al.
Published: (2025)
by: Gim, In, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
by: Lee, Junseok, et al.
Published: (2026)
by: Lee, Junseok, et al.
Published: (2026)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025)
by: Devoto, Alessio, et al.
Published: (2025)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
zFLoRA: Zero-Latency Fused Low-Rank Adapters
by: Gowda, Dhananjaya, et al.
Published: (2025)
by: Gowda, Dhananjaya, et al.
Published: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
by: Shi, Zhiyuan, et al.
Published: (2026)
by: Shi, Zhiyuan, et al.
Published: (2026)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
How Do Language Models Compose Functions?
by: Khandelwal, Apoorv, et al.
Published: (2025)
by: Khandelwal, Apoorv, et al.
Published: (2025)
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
by: Li, Ruixiao, et al.
Published: (2025)
by: Li, Ruixiao, et al.
Published: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
by: Mai, Tho, et al.
Published: (2026)
by: Mai, Tho, et al.
Published: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
by: Kim, Jongho, et al.
Published: (2025)
by: Kim, Jongho, et al.
Published: (2025)
CORM: Cache Optimization with Recent Message for Large Language Model Inference
by: Dai, Jincheng, et al.
Published: (2024)
by: Dai, Jincheng, et al.
Published: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
by: Chen, Lizhe, et al.
Published: (2025)
by: Chen, Lizhe, et al.
Published: (2025)
Cross-Attention Speculative Decoding
by: Zhong, Wei, et al.
Published: (2025)
by: Zhong, Wei, et al.
Published: (2025)
Spectral Attention Steering for Prompt Highlighting
by: Li, Weixian Waylon, et al.
Published: (2026)
by: Li, Weixian Waylon, et al.
Published: (2026)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?
by: Shah, Syed Huma
Published: (2026)
by: Shah, Syed Huma
Published: (2026)
Similar Items
-
Asynchronous LLM Function Calling
by: Gim, In, et al.
Published: (2024) -
Cacheback: Speculative Decoding With Nothing But Cache
by: Ma, Zhiyao, et al.
Published: (2025) -
Pie: A Programmable Serving System for Emerging LLM Applications
by: Gim, In, et al.
Published: (2025) -
LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
by: Bansal, Harsh Vardhan
Published: (2025) -
TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL
by: Su, Jinbo, et al.
Published: (2026)