A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Razavi, Kamran, Salmani, Mehran, Mühlhäuser, Max, Koldehofe, Boris, Wang, Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
NetNN: Neural Intrusion Detection System in Programmable Networks
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
PDSP-Bench: A Benchmarking System for Parallel and Distributed Stream Processing
von: Agnihotri, Pratyush, et al.
Veröffentlicht: (2025)
von: Agnihotri, Pratyush, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
von: Guo, Xianwen, et al.
Veröffentlicht: (2025)
von: Guo, Xianwen, et al.
Veröffentlicht: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
von: Shi, Ge, et al.
Veröffentlicht: (2025)
von: Shi, Ge, et al.
Veröffentlicht: (2025)
EcoServe: Designing Carbon-Aware AI Inference Systems
von: Li, Yueying, et al.
Veröffentlicht: (2025)
von: Li, Yueying, et al.
Veröffentlicht: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
PolyServe: Efficient Multi-SLO Serving at Scale
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
von: Zhu, Kan, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
von: Sandholm, Thomas, et al.
Veröffentlicht: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
von: Bian, Zhuohang, et al.
Veröffentlicht: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
von: Basit, Omar, et al.
Veröffentlicht: (2026)
von: Basit, Omar, et al.
Veröffentlicht: (2026)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
von: Bu, Tianci, et al.
Veröffentlicht: (2026)
von: Bu, Tianci, et al.
Veröffentlicht: (2026)
It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
von: Wu, Jing, et al.
Veröffentlicht: (2025)
von: Wu, Jing, et al.
Veröffentlicht: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
von: Wang, Weiye, et al.
Veröffentlicht: (2026)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
von: Yousefijamarani, Zahra, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
A Tale of Two Paths: Toward a Hybrid Data Plane for Efficient Far-Memory Applications
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
von: Razavi, Kamran, et al.
Veröffentlicht: (2024) -
NetNN: Neural Intrusion Detection System in Programmable Networks
von: Razavi, Kamran, et al.
Veröffentlicht: (2024) -
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023) -
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024) -
PDSP-Bench: A Benchmarking System for Parallel and Distributed Stream Processing
von: Agnihotri, Pratyush, et al.
Veröffentlicht: (2025)