PiKV: KV Cache Management System for Mixture of Experts
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Dong, Yu, Yanxuan, Lengerich, Ben, Wu, Ying Nian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)
di: Liu, Lian, et al.
Pubblicazione: (2026)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
di: Kumar, Deepak, et al.
Pubblicazione: (2025)
di: Kumar, Deepak, et al.
Pubblicazione: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
di: Canakci, Burcu, et al.
Pubblicazione: (2025)
di: Canakci, Burcu, et al.
Pubblicazione: (2025)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
Investigating Memory Failure Prediction Across CPU Architectures
di: Yu, Qiao, et al.
Pubblicazione: (2024)
di: Yu, Qiao, et al.
Pubblicazione: (2024)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
NPU Design for Diffusion Language Model Inference
di: Lou, Binglei, et al.
Pubblicazione: (2026)
di: Lou, Binglei, et al.
Pubblicazione: (2026)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2022)
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2022)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
di: Li, Jonathan, et al.
Pubblicazione: (2025)
di: Li, Jonathan, et al.
Pubblicazione: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025)
di: Choukse, Esha, et al.
Pubblicazione: (2025)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
di: Zhao, Dan, et al.
Pubblicazione: (2024)
di: Zhao, Dan, et al.
Pubblicazione: (2024)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
di: Lu, Chien-Ping
Pubblicazione: (2026)
di: Lu, Chien-Ping
Pubblicazione: (2026)
Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator
di: Hokenmaier, W, et al.
Pubblicazione: (2024)
di: Hokenmaier, W, et al.
Pubblicazione: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
di: Malenza, Giulio, et al.
Pubblicazione: (2025)
di: Malenza, Giulio, et al.
Pubblicazione: (2025)
Strict Partitioning for Sporadic Rigid Gang Tasks
di: Sun, Binqi, et al.
Pubblicazione: (2024)
di: Sun, Binqi, et al.
Pubblicazione: (2024)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
di: Dorairaj, Nij, et al.
Pubblicazione: (2026)
di: Dorairaj, Nij, et al.
Pubblicazione: (2026)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
di: Peccia, Federico Nicolas, et al.
Pubblicazione: (2024)
di: Peccia, Federico Nicolas, et al.
Pubblicazione: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Improving AI Efficiency in Data Centres by Power Dynamic Response
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
PhD Thesis Summary: Methods for Reliability Assessment and Enhancement of Deep Neural Network Hardware Accelerators
di: Taheri, Mahdi
Pubblicazione: (2026)
di: Taheri, Mahdi
Pubblicazione: (2026)
Documenti analoghi
-
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025) -
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026) -
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026) -
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
di: Zheng, Xianzhe, et al.
Pubblicazione: (2026) -
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)