Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ruihao, Pal, Shagnik, Pullu, Vineeth Narayan, Sinha, Prasoon, Ryoo, Jeeho, John, Lizy K., Yadwadkar, Neeraja J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
iServe: An Intent-based Serving System for LLMs
by: Liakopoulos, Dimitrios, et al.
Published: (2025)
by: Liakopoulos, Dimitrios, et al.
Published: (2025)
Old is Gold: Optimizing Single-threaded Applications with Exgen-Malloc
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)
by: Cai, Tianhao, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Shabari: Delayed Decision-Making for Faster and Efficient Serverless Functions
by: Sinha, Prasoon, et al.
Published: (2024)
by: Sinha, Prasoon, et al.
Published: (2024)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
by: Li, Ruihao, et al.
Published: (2026)
by: Li, Ruihao, et al.
Published: (2026)
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
by: Li, Hanchen, et al.
Published: (2025)
by: Li, Hanchen, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
Cache is King: Smart Page Eviction with eBPF
by: Zussman, Tal, et al.
Published: (2025)
by: Zussman, Tal, et al.
Published: (2025)
HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
by: Liu, Jialin, et al.
Published: (2026)
by: Liu, Jialin, et al.
Published: (2026)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
by: Nouri, Azam
Published: (2026)
by: Nouri, Azam
Published: (2026)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
by: Berend, Daniel, et al.
Published: (2025)
by: Berend, Daniel, et al.
Published: (2025)
Idiosyncrasies of Programmable Caching Engines
by: Peixoto, José, et al.
Published: (2026)
by: Peixoto, José, et al.
Published: (2026)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Collaborative dynamics in open source software development: Unveiling the influence of team interaction and the role of project manager
by: Sukrit Pal, et al.
Published: (2024)
by: Sukrit Pal, et al.
Published: (2024)
Minimal unitary dilations for commuting contractions
by: Pal, Sourav, et al.
Published: (2022)
by: Pal, Sourav, et al.
Published: (2022)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
by: Kanellis, Konstantinos, et al.
Published: (2025)
by: Kanellis, Konstantinos, et al.
Published: (2025)
RACAM: Enhancing DRAM with Reuse-Aware Computation and Automated Mapping for ML Inference
by: Ma, Siyuan, et al.
Published: (2025)
by: Ma, Siyuan, et al.
Published: (2025)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
by: Zhu, Jianian, et al.
Published: (2025)
by: Zhu, Jianian, et al.
Published: (2025)
Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU
by: Sun, He, et al.
Published: (2025)
by: Sun, He, et al.
Published: (2025)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
DPC: A Distributed Page Cache over CXL
by: Bergman, Shai, et al.
Published: (2026)
by: Bergman, Shai, et al.
Published: (2026)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Similar Items
-
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026) -
iServe: An Intent-based Serving System for LLMs
by: Liakopoulos, Dimitrios, et al.
Published: (2025) -
Old is Gold: Optimizing Single-threaded Applications with Exgen-Malloc
by: Li, Ruihao, et al.
Published: (2025) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025) -
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)