Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Razavi, Kamran, Ghafouri, Saeid, Mühlhäuser, Max, Jamshidi, Pooyan, Wang, Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
di: Razavi, Kamran, et al.
Pubblicazione: (2024)
di: Razavi, Kamran, et al.
Pubblicazione: (2024)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
di: Ghafouri, Saeid, et al.
Pubblicazione: (2023)
di: Ghafouri, Saeid, et al.
Pubblicazione: (2023)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
di: Chen, Siyuan, et al.
Pubblicazione: (2025)
di: Chen, Siyuan, et al.
Pubblicazione: (2025)
NetNN: Neural Intrusion Detection System in Programmable Networks
di: Razavi, Kamran, et al.
Pubblicazione: (2024)
di: Razavi, Kamran, et al.
Pubblicazione: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
di: Santavas, Nicholas, et al.
Pubblicazione: (2026)
di: Santavas, Nicholas, et al.
Pubblicazione: (2026)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
di: Wang, Weiye, et al.
Pubblicazione: (2026)
di: Wang, Weiye, et al.
Pubblicazione: (2026)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
di: Li, Yueying, et al.
Pubblicazione: (2025)
di: Li, Yueying, et al.
Pubblicazione: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
di: Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
di: Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
di: Li, Xiangchen, et al.
Pubblicazione: (2025)
di: Li, Xiangchen, et al.
Pubblicazione: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
di: Lin, Yi-Chien, et al.
Pubblicazione: (2024)
Enabling Dynamic Sparsity in Quantized LLM Inference
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
di: Liyanage, Mohan, et al.
Pubblicazione: (2026)
di: Liyanage, Mohan, et al.
Pubblicazione: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
di: Shi, Ge, et al.
Pubblicazione: (2025)
di: Shi, Ge, et al.
Pubblicazione: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
di: Huang, Heyang, et al.
Pubblicazione: (2025)
di: Huang, Heyang, et al.
Pubblicazione: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
di: Wolfrath, Joel, et al.
Pubblicazione: (2025)
di: Wolfrath, Joel, et al.
Pubblicazione: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
di: Guo, Xianwen, et al.
Pubblicazione: (2025)
di: Guo, Xianwen, et al.
Pubblicazione: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
di: Kossmann, Ferdi, et al.
Pubblicazione: (2024)
di: Kossmann, Ferdi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
di: Razavi, Kamran, et al.
Pubblicazione: (2024) -
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
di: Ghafouri, Saeid, et al.
Pubblicazione: (2023) -
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
di: Chen, Siyuan, et al.
Pubblicazione: (2025) -
NetNN: Neural Intrusion Detection System in Programmable Networks
di: Razavi, Kamran, et al.
Pubblicazione: (2024) -
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)