PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sun, Desen, Zhao, Zepeng, Wang, Yuke |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
par: Ye, Fanjiang, et autres
Publié: (2026)
par: Ye, Fanjiang, et autres
Publié: (2026)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025)
par: Chen, Siyuan, et autres
Publié: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024)
par: Shen, Haiying, et autres
Publié: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
par: Tian, Yuyang, et autres
Publié: (2025)
par: Tian, Yuyang, et autres
Publié: (2025)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
par: Yousefijamarani, Zahra, et autres
Publié: (2025)
par: Yousefijamarani, Zahra, et autres
Publié: (2025)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
par: Yang, Qizheng, et autres
Publié: (2025)
par: Yang, Qizheng, et autres
Publié: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
par: Li, Zikun, et autres
Publié: (2025)
par: Li, Zikun, et autres
Publié: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
par: Ma, Chenxiang, et autres
Publié: (2025)
par: Ma, Chenxiang, et autres
Publié: (2025)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
par: Gao, Wei, et autres
Publié: (2026)
par: Gao, Wei, et autres
Publié: (2026)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
par: Sedlak, Boris, et autres
Publié: (2024)
par: Sedlak, Boris, et autres
Publié: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
par: Xia, Yifei, et autres
Publié: (2025)
par: Xia, Yifei, et autres
Publié: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
par: Huang, Heyang, et autres
Publié: (2025)
par: Huang, Heyang, et autres
Publié: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
par: Huang, Shaoyuan, et autres
Publié: (2026)
par: Huang, Shaoyuan, et autres
Publié: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
par: Qi, S., et autres
Publié: (2024)
par: Qi, S., et autres
Publié: (2024)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
par: Zhang, Zongshun, et autres
Publié: (2025)
par: Zhang, Zongshun, et autres
Publié: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
par: Xuan, Mo, et autres
Publié: (2025)
par: Xuan, Mo, et autres
Publié: (2025)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
par: Punniyamoorthy, Vinoth, et autres
Publié: (2025)
JITServe: SLO-aware LLM Serving with Imprecise Request Information
par: Zhang, Wei, et autres
Publié: (2025)
par: Zhang, Wei, et autres
Publié: (2025)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
par: Wu, Zhiyuan, et autres
Publié: (2025)
par: Wu, Zhiyuan, et autres
Publié: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted Microservices
par: Wang, Zibo, et autres
Publié: (2022)
par: Wang, Zibo, et autres
Publié: (2022)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
par: Jeon, Beomyeol, et autres
Publié: (2024)
par: Jeon, Beomyeol, et autres
Publié: (2024)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
par: Chen, Jiabin, et autres
Publié: (2024)
par: Chen, Jiabin, et autres
Publié: (2024)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
par: Peng, Haosong, et autres
Publié: (2024)
par: Peng, Haosong, et autres
Publié: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
par: Jaiswal, Shashwat, et autres
Publié: (2025)
par: Jaiswal, Shashwat, et autres
Publié: (2025)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
par: Oliaro, Gabriele, et autres
Publié: (2024)
par: Oliaro, Gabriele, et autres
Publié: (2024)
Documents similaires
-
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
par: Ye, Fanjiang, et autres
Publié: (2026) -
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025) -
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026) -
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024) -
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)