KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yichun, Khaira, Navjot K., Singh, Tejinder |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is (Selective) Round-To-Nearest Quantization All You Need?
von: Kogan, Alex
Veröffentlicht: (2025)
von: Kogan, Alex
Veröffentlicht: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
von: Sirotkina, Elena
Veröffentlicht: (2026)
von: Sirotkina, Elena
Veröffentlicht: (2026)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations
von: Karlsson, Robin, et al.
Veröffentlicht: (2026)
von: Karlsson, Robin, et al.
Veröffentlicht: (2026)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
von: Zolduoarrati, Elijah, et al.
Veröffentlicht: (2025)
von: Zolduoarrati, Elijah, et al.
Veröffentlicht: (2025)
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
von: Dong, Wenxin, et al.
Veröffentlicht: (2026)
von: Dong, Wenxin, et al.
Veröffentlicht: (2026)
Secure and Scalable Blockchain Voting: A Comparative Framework and the Role of Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
von: Nemitz, Jonathan, et al.
Veröffentlicht: (2026)
von: Nemitz, Jonathan, et al.
Veröffentlicht: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
Neural Attention: A Novel Mechanism for Enhanced Expressive Power in Transformer Models
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
von: DiGiugno, Andrew, et al.
Veröffentlicht: (2025)
Development of a Smart Autonomous Irrigation System Using Iot and AI
von: Kunt, Yunus Emre
Veröffentlicht: (2025)
von: Kunt, Yunus Emre
Veröffentlicht: (2025)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
von: Rudman, William, et al.
Veröffentlicht: (2026)
von: Rudman, William, et al.
Veröffentlicht: (2026)
Using Deep Learning to Generate Semantically Correct Hindi Captions
von: Khan, Wasim Akram, et al.
Veröffentlicht: (2026)
von: Khan, Wasim Akram, et al.
Veröffentlicht: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
von: Abualazm, Raafat, et al.
Veröffentlicht: (2026)
von: Abualazm, Raafat, et al.
Veröffentlicht: (2026)
When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
von: Li, Haorui, et al.
Veröffentlicht: (2026)
von: Li, Haorui, et al.
Veröffentlicht: (2026)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
von: Othman, Refat
Veröffentlicht: (2026)
von: Othman, Refat
Veröffentlicht: (2026)
AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
von: Loi, Dario, et al.
Veröffentlicht: (2025)
von: Loi, Dario, et al.
Veröffentlicht: (2025)
Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models
von: Wang, Junxin, et al.
Veröffentlicht: (2026)
von: Wang, Junxin, et al.
Veröffentlicht: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
von: Lucas, Tom, et al.
Veröffentlicht: (2026)
A Faster and More Reliable Middleware for Autonomous Driving Systems
von: He, Yuankai, et al.
Veröffentlicht: (2025)
von: He, Yuankai, et al.
Veröffentlicht: (2025)
Supervised Embedded Methods for Hyperspectral Band Selection
von: Zimmer, Yaniv, et al.
Veröffentlicht: (2024)
von: Zimmer, Yaniv, et al.
Veröffentlicht: (2024)
NetCAS: Dynamic Cache and Backend Device Management in Networked Environments
von: Hwang, Joon Yong, et al.
Veröffentlicht: (2025)
von: Hwang, Joon Yong, et al.
Veröffentlicht: (2025)
Framework Matters: Energy Efficiency of UI Automation Testing Frameworks
von: Lagermann, Timmie M. R., et al.
Veröffentlicht: (2025)
von: Lagermann, Timmie M. R., et al.
Veröffentlicht: (2025)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
von: Pandey, Anupam, et al.
Veröffentlicht: (2025)
Secure coding for web applications: Frameworks, challenges, and the role of LLMs
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Is (Selective) Round-To-Nearest Quantization All You Need?
von: Kogan, Alex
Veröffentlicht: (2025) -
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
von: Liu, Zhuoyao, et al.
Veröffentlicht: (2026) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026) -
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026) -
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
von: Sirotkina, Elena
Veröffentlicht: (2026)