GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Boxiao, Huangfu, Boning, Luo, Yizhou, Chen, Chen, Li, Zijun, Yu, Minchen, Fan, Xiaoyi, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
von: He, Xuan, et al.
Veröffentlicht: (2025)
von: He, Xuan, et al.
Veröffentlicht: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
von: Shen, Haiying, et al.
Veröffentlicht: (2024)
Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jinyuan, et al.
Veröffentlicht: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
von: Lin, Zejia, et al.
Veröffentlicht: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
von: Pagonas, Nikos, et al.
Veröffentlicht: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
von: Liu, Yifei, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
Fast State Restoration in LLM Serving with HCache
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
von: Dai, Yinwei, et al.
Veröffentlicht: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
von: Zhao, Zhixin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixin, et al.
Veröffentlicht: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024) -
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025) -
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024) -
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025) -
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)