Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
Fuente:
arXiv
Saved in:
| Main Author: | Ganjihal, Sanjeev Rao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026)
by: Jo, Myeong Jun
Published: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026)
by: Georgiou, Athos
Published: (2026)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
by: Zhang, Yongkang, et al.
Published: (2024)
by: Zhang, Yongkang, et al.
Published: (2024)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
by: Lo, Yun-Chen, et al.
Published: (2024)
by: Lo, Yun-Chen, et al.
Published: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
by: Chen, Mu-Chi, et al.
Published: (2025)
by: Chen, Mu-Chi, et al.
Published: (2025)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026)
by: Topcu, Burak, et al.
Published: (2026)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
by: Kamath, Aditya K, et al.
Published: (2026)
by: Kamath, Aditya K, et al.
Published: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)
by: Zheng, Xianzhe, et al.
Published: (2026)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
by: Tran, Brandon, et al.
Published: (2026)
by: Tran, Brandon, et al.
Published: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference
by: Nguyen, An Xuan
Published: (2026)
by: Nguyen, An Xuan
Published: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
by: Adhinarayanan, Vignesh, et al.
Published: (2026)
Random Adaptive Cache Placement Policy
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
UPMEM Unleashed: Software Secrets for Speed
by: Chmielewski, Krystian, et al.
Published: (2025)
by: Chmielewski, Krystian, et al.
Published: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Can Asymmetric Tile Buffering Be Beneficial?
by: Wang, Chengyue, et al.
Published: (2025)
by: Wang, Chengyue, et al.
Published: (2025)
Exploiting long vectors with a CFD code: a co-design show case
by: Blancafort, Marc, et al.
Published: (2024)
by: Blancafort, Marc, et al.
Published: (2024)
Similar Items
-
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
by: Jo, Myeong Jun
Published: (2026) -
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
by: Georgiou, Athos
Published: (2026) -
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
by: Zhang, Yongkang, et al.
Published: (2024) -
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
by: Lo, Yun-Chen, et al.
Published: (2024) -
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)