GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ye, Fanjiang, Li, Zhangke, Zhong, Xinrui, Ma, Ethan, Chen, Russell, Wang, Kaijian, Zuo, Jingwei, Sun, Desen, Cao, Ye, Cao, Triston, Lee, Myungjin, Krishnamurthy, Arvind, Wang, Yuke |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
par: Zuo, Jingwei, et autres
Publié: (2026)
par: Zuo, Jingwei, et autres
Publié: (2026)
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
par: Zhong, Xinrui, et autres
Publié: (2025)
par: Zhong, Xinrui, et autres
Publié: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
par: Sun, Desen, et autres
Publié: (2025)
par: Sun, Desen, et autres
Publié: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
par: Pagonas, Nikos, et autres
Publié: (2025)
par: Pagonas, Nikos, et autres
Publié: (2025)
Enabling Elastic Model Serving with MultiWorld
par: Lee, Myungjin, et autres
Publié: (2024)
par: Lee, Myungjin, et autres
Publié: (2024)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025)
par: Chen, Siyuan, et autres
Publié: (2025)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
par: Zhu, Kan, et autres
Publié: (2024)
par: Zhu, Kan, et autres
Publié: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
par: Wilkins, Grant, et autres
Publié: (2024)
par: Wilkins, Grant, et autres
Publié: (2024)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
par: Jiang, Youhe, et autres
Publié: (2026)
par: Jiang, Youhe, et autres
Publié: (2026)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
par: Xiang, Yuxing, et autres
Publié: (2025)
par: Xiang, Yuxing, et autres
Publié: (2025)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
par: Liu, Rixin, et autres
Publié: (2026)
par: Liu, Rixin, et autres
Publié: (2026)
Software-Defined Agentic Serving
par: Agarwal, Saurabh, et autres
Publié: (2026)
par: Agarwal, Saurabh, et autres
Publié: (2026)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
par: Luo, Ziyue, et autres
Publié: (2025)
par: Luo, Ziyue, et autres
Publié: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
par: Wang, Zhigang, et autres
Publié: (2024)
par: Wang, Zhigang, et autres
Publié: (2024)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
par: Tian, Yuyang, et autres
Publié: (2025)
par: Tian, Yuyang, et autres
Publié: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
par: Go, Seokjin, et autres
Publié: (2026)
par: Go, Seokjin, et autres
Publié: (2026)
Efficient Unified Caching for Accelerating Heterogeneous AI Workloads
par: Wang, Tianze, et autres
Publié: (2025)
par: Wang, Tianze, et autres
Publié: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
par: Wang, Yuxin, et autres
Publié: (2024)
par: Wang, Yuxin, et autres
Publié: (2024)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
par: Chen, Lequn, et autres
Publié: (2023)
par: Chen, Lequn, et autres
Publié: (2023)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
par: Hu, Bodun, et autres
Publié: (2024)
par: Hu, Bodun, et autres
Publié: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
par: Cao, Jiahe, et autres
Publié: (2026)
par: Cao, Jiahe, et autres
Publié: (2026)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
par: Su, Zhaoyuan, et autres
Publié: (2025)
par: Su, Zhaoyuan, et autres
Publié: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
par: Zhou, Qihui, et autres
Publié: (2025)
par: Zhou, Qihui, et autres
Publié: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
par: Tang, Lujie, et autres
Publié: (2024)
par: Tang, Lujie, et autres
Publié: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
par: Du, Boxiao, et autres
Publié: (2026)
par: Du, Boxiao, et autres
Publié: (2026)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
par: Wen, Linfeng, et autres
Publié: (2024)
par: Wen, Linfeng, et autres
Publié: (2024)
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
par: Zhang, Boyuan, et autres
Publié: (2024)
par: Zhang, Boyuan, et autres
Publié: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
par: Zhao, Bohan, et autres
Publié: (2025)
par: Zhao, Bohan, et autres
Publié: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
par: Chung, Jae-Won, et autres
Publié: (2026)
par: Chung, Jae-Won, et autres
Publié: (2026)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
par: Wang, Irene, et autres
Publié: (2026)
par: Wang, Irene, et autres
Publié: (2026)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
par: Wang, Chao, et autres
Publié: (2025)
par: Wang, Chao, et autres
Publié: (2025)
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
par: Yuan, Ying, et autres
Publié: (2026)
par: Yuan, Ying, et autres
Publié: (2026)
Documents similaires
-
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
par: Zuo, Jingwei, et autres
Publié: (2026) -
An Efficient and Adaptive Watermark Detection System with Tile-based Error Correction
par: Zhong, Xinrui, et autres
Publié: (2025) -
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
par: Sun, Desen, et autres
Publié: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025) -
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
par: Pagonas, Nikos, et autres
Publié: (2025)