Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Razavi, Kamran, Ghafouri, Saeid, Mühlhäuser, Max, Jamshidi, Pooyan, Wang, Lin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
par: Razavi, Kamran, et autres
Publié: (2024)
par: Razavi, Kamran, et autres
Publié: (2024)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
par: Ghafouri, Saeid, et autres
Publié: (2023)
par: Ghafouri, Saeid, et autres
Publié: (2023)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025)
par: Chen, Siyuan, et autres
Publié: (2025)
NetNN: Neural Intrusion Detection System in Programmable Networks
par: Razavi, Kamran, et autres
Publié: (2024)
par: Razavi, Kamran, et autres
Publié: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025)
par: Hu, Junhao, et autres
Publié: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
par: Li, Suyi, et autres
Publié: (2024)
par: Li, Suyi, et autres
Publié: (2024)
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
par: Santavas, Nicholas, et autres
Publié: (2026)
par: Santavas, Nicholas, et autres
Publié: (2026)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
Serving Compound Inference Systems on Datacenter GPUs
par: Devata, Sriram, et autres
Publié: (2026)
par: Devata, Sriram, et autres
Publié: (2026)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025)
par: Wu, Jingfeng, et autres
Publié: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
par: Du, Boxiao, et autres
Publié: (2026)
par: Du, Boxiao, et autres
Publié: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
par: Jaiswal, Shashwat, et autres
Publié: (2025)
par: Jaiswal, Shashwat, et autres
Publié: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
par: Wang, Weiye, et autres
Publié: (2026)
par: Wang, Weiye, et autres
Publié: (2026)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
par: Li, Yueying, et autres
Publié: (2025)
par: Li, Yueying, et autres
Publié: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
par: Nguyen, Thanh-Tung, et autres
Publié: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
par: Li, Xiangchen, et autres
Publié: (2025)
par: Li, Xiangchen, et autres
Publié: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
par: Chen, Wenyan, et autres
Publié: (2026)
par: Chen, Wenyan, et autres
Publié: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
par: Zhou, Qihui, et autres
Publié: (2025)
par: Zhou, Qihui, et autres
Publié: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
par: Bai, Fan, et autres
Publié: (2026)
par: Bai, Fan, et autres
Publié: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
par: Lin, Yi-Chien, et autres
Publié: (2024)
par: Lin, Yi-Chien, et autres
Publié: (2024)
Enabling Dynamic Sparsity in Quantized LLM Inference
par: Wang, Rongxiang, et autres
Publié: (2025)
par: Wang, Rongxiang, et autres
Publié: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
Risk-Aware and Stable Edge Server Selection Under Network Latency SLOs
par: Liyanage, Mohan, et autres
Publié: (2026)
par: Liyanage, Mohan, et autres
Publié: (2026)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
par: Chen, Xing, et autres
Publié: (2025)
par: Chen, Xing, et autres
Publié: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
par: Nie, Chengyi, et autres
Publié: (2024)
par: Nie, Chengyi, et autres
Publié: (2024)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
par: Shi, Ge, et autres
Publié: (2025)
par: Shi, Ge, et autres
Publié: (2025)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
par: Huang, Heyang, et autres
Publié: (2025)
par: Huang, Heyang, et autres
Publié: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
par: Wolfrath, Joel, et autres
Publié: (2025)
par: Wolfrath, Joel, et autres
Publié: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
par: Guo, Xianwen, et autres
Publié: (2025)
par: Guo, Xianwen, et autres
Publié: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
CascadeServe: Unlocking Model Cascades for Inference Serving
par: Kossmann, Ferdi, et autres
Publié: (2024)
par: Kossmann, Ferdi, et autres
Publié: (2024)
Documents similaires
-
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
par: Razavi, Kamran, et autres
Publié: (2024) -
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
par: Ghafouri, Saeid, et autres
Publié: (2023) -
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
par: Chen, Siyuan, et autres
Publié: (2025) -
NetNN: Neural Intrusion Detection System in Programmable Networks
par: Razavi, Kamran, et autres
Publié: (2024) -
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)