Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Haoyu, Fu, Fangcheng, Wu, Jia, Yuan, Binhang, Zhang, Yongqiang, Wang, Hao, Zhu, Yuanyuan, Yan, Xiao, Jiang, Jiawei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
Cascadia: An Efficient Cascade Serving System for Large Language Models
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026)
di: Peng, You, et al.
Pubblicazione: (2026)
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024)
di: Xiong, Yi, et al.
Pubblicazione: (2024)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
di: Guo, Tianyu, et al.
Pubblicazione: (2025)
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
di: Zhong, Zhiqing, et al.
Pubblicazione: (2026)
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration
di: Nian, Sean, et al.
Pubblicazione: (2026)
di: Nian, Sean, et al.
Pubblicazione: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
di: Yan, Ran, et al.
Pubblicazione: (2024)
di: Yan, Ran, et al.
Pubblicazione: (2024)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
di: Xiang, Xingyu, et al.
Pubblicazione: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
di: Qiu, Shi, et al.
Pubblicazione: (2026)
di: Qiu, Shi, et al.
Pubblicazione: (2026)
Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
di: Wang, Yuhang, et al.
Pubblicazione: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
di: Qianli, Liu, et al.
Pubblicazione: (2025)
di: Qianli, Liu, et al.
Pubblicazione: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
di: Wang, Yixuan, et al.
Pubblicazione: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
di: Liu, Zedong, et al.
Pubblicazione: (2026)
di: Liu, Zedong, et al.
Pubblicazione: (2026)
MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches
di: Wang, Xin, et al.
Pubblicazione: (2026)
di: Wang, Xin, et al.
Pubblicazione: (2026)
Pythia: Exploiting Workflow Predictability for Efficient Agent-Native LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2026)
di: Yu, Shan, et al.
Pubblicazione: (2026)
Cost-Efficient LLM Serving in the Cloud: VM Selection with KV Cache Offloading
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
di: Kim, Kihyun, et al.
Pubblicazione: (2025)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
di: Wang, Shao, et al.
Pubblicazione: (2026)
di: Wang, Shao, et al.
Pubblicazione: (2026)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
di: Liu, Yuhan, et al.
Pubblicazione: (2023)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
di: Bian, Zhuohang, et al.
Pubblicazione: (2025)
di: Bian, Zhuohang, et al.
Pubblicazione: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
di: li, Fei, et al.
Pubblicazione: (2026)
di: li, Fei, et al.
Pubblicazione: (2026)
FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
di: Zhu, Jianian, et al.
Pubblicazione: (2025)
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention
di: Xu, Dongjie, et al.
Pubblicazione: (2026)
di: Xu, Dongjie, et al.
Pubblicazione: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: Yiyuan He, et al.
Pubblicazione: (2026)
di: Yiyuan He, et al.
Pubblicazione: (2026)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
di: Feng, Yuan, et al.
Pubblicazione: (2025)
di: Feng, Yuan, et al.
Pubblicazione: (2025)
Joint Encoding of KV-Cache Blocks for Scalable LLM Serving
di: Kampeas, Joseph, et al.
Pubblicazione: (2026)
di: Kampeas, Joseph, et al.
Pubblicazione: (2026)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
di: Xia, Yifei, et al.
Pubblicazione: (2025)
di: Xia, Yifei, et al.
Pubblicazione: (2025)
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
di: Chen, Kaiwen, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
InstCache: A Predictive Cache for LLM Serving
di: Zou, Longwei, et al.
Pubblicazione: (2024)
di: Zou, Longwei, et al.
Pubblicazione: (2024)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
di: Jiang, Youhe, et al.
Pubblicazione: (2026) -
Cascadia: An Efficient Cascade Serving System for Large Language Models
di: Jiang, Youhe, et al.
Pubblicazione: (2025) -
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026) -
LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
di: Xiong, Yi, et al.
Pubblicazione: (2024) -
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)