Saved in:
| Main Authors: | Liu, Jialin, Shi, Liang, Yu, Dingcui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.01655 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
by: Yu, Dingcui, et al.
Published: (2025)
by: Yu, Dingcui, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
by: Kanellopoulos, Konstantinos, et al.
Published: (2023)
Cache is King: Smart Page Eviction with eBPF
by: Zussman, Tal, et al.
Published: (2025)
by: Zussman, Tal, et al.
Published: (2025)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
by: Berend, Daniel, et al.
Published: (2025)
by: Berend, Daniel, et al.
Published: (2025)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
by: Li, Haoyu, et al.
Published: (2025)
by: Li, Haoyu, et al.
Published: (2025)
Concurrency Testing in the Linux Kernel via eBPF
by: Xu, Jiacheng, et al.
Published: (2025)
by: Xu, Jiacheng, et al.
Published: (2025)
Leveraging OS-Level Primitives for Robotic Action Management
by: Zheng, Wenxin, et al.
Published: (2025)
by: Zheng, Wenxin, et al.
Published: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)
by: Zhao, Haoru, et al.
Published: (2025)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
by: Cai, Tianhao, et al.
Published: (2025)
by: Cai, Tianhao, et al.
Published: (2025)
Idiosyncrasies of Programmable Caching Engines
by: Peixoto, José, et al.
Published: (2026)
by: Peixoto, José, et al.
Published: (2026)
Principled Performance Tunability in Operating System Kernels
by: Chen, Zhongjie, et al.
Published: (2025)
by: Chen, Zhongjie, et al.
Published: (2025)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025)
by: Han, Mingcong, et al.
Published: (2025)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
by: Smejkal, Till, et al.
Published: (2024)
by: Smejkal, Till, et al.
Published: (2024)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
Age-Memory Trade-off in Read-Copy-Update
by: Ramani, Vishakha, et al.
Published: (2024)
by: Ramani, Vishakha, et al.
Published: (2024)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
by: Zhang, Dingyan, et al.
Published: (2024)
by: Zhang, Dingyan, et al.
Published: (2024)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
RUISA Operational Ecosystem Architecture
by: AL Mohtar, Mouayad
Published: (2026)
by: AL Mohtar, Mouayad
Published: (2026)
DPC: A Distributed Page Cache over CXL
by: Bergman, Shai, et al.
Published: (2026)
by: Bergman, Shai, et al.
Published: (2026)
I/O Transit Caching for PMem-based Block Device
by: Xu, Qing, et al.
Published: (2024)
by: Xu, Qing, et al.
Published: (2024)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
by: Kanellis, Konstantinos, et al.
Published: (2025)
by: Kanellis, Konstantinos, et al.
Published: (2025)
Assessing FIFO and Round Robin Scheduling:Effects on Data Pipeline Performance and Energy Usage
by: Choudhury, Malobika Roy, et al.
Published: (2024)
by: Choudhury, Malobika Roy, et al.
Published: (2024)
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
by: Liu, Fangyue, et al.
Published: (2026)
by: Liu, Fangyue, et al.
Published: (2026)
Phoenix -- A Novel Technique for Performance-Aware Orchestration of Thread and Page Table Placement in NUMA Systems
by: Siavashi, Mohammad, et al.
Published: (2025)
by: Siavashi, Mohammad, et al.
Published: (2025)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
The First Principle of Big Memory Systems
by: Hua, Yu
Published: (2023)
by: Hua, Yu
Published: (2023)
Qurator: Scheduling Hybrid Quantum-Classical Workflows Across Heterogeneous Cloud Providers
by: Pehlivanoglu, Sinan, et al.
Published: (2026)
by: Pehlivanoglu, Sinan, et al.
Published: (2026)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
by: Kang, Duseok, et al.
Published: (2025)
by: Kang, Duseok, et al.
Published: (2025)
Similar Items
-
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
by: Yu, Dingcui, et al.
Published: (2025) -
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026) -
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
by: Kanellopoulos, Konstantinos, et al.
Published: (2023) -
Cache is King: Smart Page Eviction with eBPF
by: Zussman, Tal, et al.
Published: (2025) -
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
by: Berend, Daniel, et al.
Published: (2025)