CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Weiye, Chen, Chen, Zhang, Junxue, Wang, Zhusheng, Yuan, Hui, Guan, Zixuan, Zheng, Xiaolong, Weng, Qizhen, Chen, Yin, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexiang, et al.
Veröffentlicht: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
von: Wang, Tianze, et al.
Veröffentlicht: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
von: Yuan, Ying, et al.
Veröffentlicht: (2026)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
von: Liu, Zihan, et al.
Veröffentlicht: (2025)
Fast State Restoration in LLM Serving with HCache
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
von: Chen, Haoyu, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
von: Fang, Jingzhi, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
von: Yang, Qizheng, et al.
Veröffentlicht: (2025)
von: Yang, Qizheng, et al.
Veröffentlicht: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
von: Wu, Siyu, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
von: Zhang, Yuning, et al.
Veröffentlicht: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024) -
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
von: Liu, Yifei, et al.
Veröffentlicht: (2025) -
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026) -
SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity
von: Gan, Zhenghao, et al.
Veröffentlicht: (2026) -
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)