Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zheng, Xianzhe, Wang, Zhengheng, Ma, Ruiyan, Wang, Rui, Wang, Xiyu, Chen, Rui, Zhang, Peng, Pan, Sicheng, Huang, Zhangheng, Wu, Chenxin, Zhang, Yi, Cai, Bo, Liu, Kan, Ma, Teng, Du, Yin, Deng, Dong, Wu, Sai, Zhu, Guoyun, Zhang, Wei, Li, Feifei |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025)
par: Fang, Yunhua, et autres
Publié: (2025)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
par: Xia, Tianhua, et autres
Publié: (2025)
par: Xia, Tianhua, et autres
Publié: (2025)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
par: Yao, Jiayi, et autres
Publié: (2026)
par: Yao, Jiayi, et autres
Publié: (2026)
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
par: Wang, Zhican, et autres
Publié: (2025)
par: Wang, Zhican, et autres
Publié: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
par: li, Fei, et autres
Publié: (2026)
par: li, Fei, et autres
Publié: (2026)
PiKV: KV Cache Management System for Mixture of Experts
par: Liu, Dong, et autres
Publié: (2025)
par: Liu, Dong, et autres
Publié: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
par: Zhang, Junming, et autres
Publié: (2026)
par: Zhang, Junming, et autres
Publié: (2026)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
par: Wang, Quancheng, et autres
Publié: (2023)
par: Wang, Quancheng, et autres
Publié: (2023)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
par: Chen, Peilin, et autres
Publié: (2025)
par: Chen, Peilin, et autres
Publié: (2025)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
par: Chen, Yiqi, et autres
Publié: (2025)
par: Chen, Yiqi, et autres
Publié: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
par: Zhou, Zhe, et autres
Publié: (2024)
par: Zhou, Zhe, et autres
Publié: (2024)
RoSE-Opt: Robust and Efficient Analog Circuit Parameter Optimization with Knowledge-infused Reinforcement Learning
par: Cao, Weidong, et autres
Publié: (2024)
par: Cao, Weidong, et autres
Publié: (2024)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
par: Mamo, Oteo, et autres
Publié: (2026)
par: Mamo, Oteo, et autres
Publié: (2026)
DCI: A Coordinated Allocation and Filling Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
par: Luo, Yi, et autres
Publié: (2025)
par: Luo, Yi, et autres
Publié: (2025)
Changing the Game: The Bounce-Bind Ising Machine
par: Zhang, Haiyang, et autres
Publié: (2026)
par: Zhang, Haiyang, et autres
Publié: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
par: Zhang, Hang, et autres
Publié: (2025)
par: Zhang, Hang, et autres
Publié: (2025)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
par: Hong, Jeongmin, et autres
Publié: (2024)
par: Hong, Jeongmin, et autres
Publié: (2024)
An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision Sensors
par: Zhao, Qinghang, et autres
Publié: (2024)
par: Zhao, Qinghang, et autres
Publié: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
par: Ganjihal, Sanjeev Rao
Publié: (2026)
par: Ganjihal, Sanjeev Rao
Publié: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
par: Kim, Minsu, et autres
Publié: (2025)
par: Kim, Minsu, et autres
Publié: (2025)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
par: Xie, Rui, et autres
Publié: (2024)
par: Xie, Rui, et autres
Publié: (2024)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
par: Xu, Weikai, et autres
Publié: (2025)
par: Xu, Weikai, et autres
Publié: (2025)
AccelSync: Verifying Synchronization Coverage in Accelerator Pipeline Programs
par: An, Hangcheng, et autres
Publié: (2026)
par: An, Hangcheng, et autres
Publié: (2026)
AutoPDR: Circuit-Aware Solver Configuration Prediction for Hardware Model Checking
par: Hu, Guangyu, et autres
Publié: (2026)
par: Hu, Guangyu, et autres
Publié: (2026)
Explainable Fuzzy Neural Network with Multi-Fidelity Reinforcement Learning for Micro-Architecture Design Space Exploration
par: Fan, Hanwei, et autres
Publié: (2024)
par: Fan, Hanwei, et autres
Publié: (2024)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
par: Zhang, Tao, et autres
Publié: (2026)
par: Zhang, Tao, et autres
Publié: (2026)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
par: He, Zifan, et autres
Publié: (2025)
par: He, Zifan, et autres
Publié: (2025)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
par: Yamamoto, Yuji, et autres
Publié: (2026)
par: Yamamoto, Yuji, et autres
Publié: (2026)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
par: Zhao, Wei, et autres
Publié: (2024)
par: Zhao, Wei, et autres
Publié: (2024)
RHS-TRNG: A Resilient High-Speed True Random Number Generator Based on STT-MTJ Device
par: Fu, Siqing, et autres
Publié: (2023)
par: Fu, Siqing, et autres
Publié: (2023)
Hyft: A Reconfigurable Softmax Accelerator with Hybrid Numeric Format for both Training and Inference
par: Xia, Tianhua, et autres
Publié: (2023)
par: Xia, Tianhua, et autres
Publié: (2023)
Architectural and System Implications of CXL-enabled Tiered Memory
par: Yang, Yujie, et autres
Publié: (2025)
par: Yang, Yujie, et autres
Publié: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
par: You, Dean, et autres
Publié: (2025)
par: You, Dean, et autres
Publié: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
A Near-Cache Architectural Framework for Cryptographic Computing
par: Zhang, Jingyao, et autres
Publié: (2025)
par: Zhang, Jingyao, et autres
Publié: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
par: Wang, Peipei, et autres
Publié: (2025)
par: Wang, Peipei, et autres
Publié: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
par: Wang, Haoxuan, et autres
Publié: (2026)
par: Wang, Haoxuan, et autres
Publié: (2026)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
par: Huang, Zhirui, et autres
Publié: (2025)
par: Huang, Zhirui, et autres
Publié: (2025)
PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
par: Li, Luyi, et autres
Publié: (2023)
par: Li, Luyi, et autres
Publié: (2023)
Documents similaires
-
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025) -
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
par: Xia, Tianhua, et autres
Publié: (2025) -
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
par: Yao, Jiayi, et autres
Publié: (2026) -
VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
par: Wang, Zhican, et autres
Publié: (2025) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
par: li, Fei, et autres
Publié: (2026)