Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahmad, Sohaib, Guan, Hui, Sitaraman, Ramesh K. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
por: Ahmad, Sohaib, et al.
Publicado: (2024)
por: Ahmad, Sohaib, et al.
Publicado: (2024)
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
por: Yang, Qizheng, et al.
Publicado: (2025)
por: Yang, Qizheng, et al.
Publicado: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
por: Xia, Yifei, et al.
Publicado: (2025)
por: Xia, Yifei, et al.
Publicado: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
por: Razavi, Kamran, et al.
Publicado: (2024)
por: Razavi, Kamran, et al.
Publicado: (2024)
Declarative Data Pipeline for Large Scale ML Services
por: Yang, Yunzhao, et al.
Publicado: (2025)
por: Yang, Yunzhao, et al.
Publicado: (2025)
The Green Mirage: Impact of Location- and Market-based Carbon Intensity Estimation on Carbon Optimization Efficacy
por: Maji, Diptyaroop, et al.
Publicado: (2024)
por: Maji, Diptyaroop, et al.
Publicado: (2024)
Untangling Carbon-free Energy Attribution and Carbon Intensity Estimation for Carbon-aware Computing
por: Maji, Diptyaroop, et al.
Publicado: (2023)
por: Maji, Diptyaroop, et al.
Publicado: (2023)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
por: Wang, Weiye, et al.
Publicado: (2026)
por: Wang, Weiye, et al.
Publicado: (2026)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
por: Razavi, Kamran, et al.
Publicado: (2024)
por: Razavi, Kamran, et al.
Publicado: (2024)
EcoServe: Designing Carbon-Aware AI Inference Systems
por: Li, Yueying, et al.
Publicado: (2025)
por: Li, Yueying, et al.
Publicado: (2025)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
por: Guo, Tianyu, et al.
Publicado: (2025)
por: Guo, Tianyu, et al.
Publicado: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
por: Hu, Junhao, et al.
Publicado: (2025)
por: Hu, Junhao, et al.
Publicado: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
por: Guo, Xianwen, et al.
Publicado: (2025)
por: Guo, Xianwen, et al.
Publicado: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
por: Zhan, Huiyou, et al.
Publicado: (2025)
por: Zhan, Huiyou, et al.
Publicado: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
por: Li, Suyi, et al.
Publicado: (2024)
por: Li, Suyi, et al.
Publicado: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
por: He, Yiyuan, et al.
Publicado: (2024)
por: He, Yiyuan, et al.
Publicado: (2024)
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
por: Yoshimura, Takeshi, et al.
Publicado: (2026)
por: Yoshimura, Takeshi, et al.
Publicado: (2026)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
por: Hu, Jianmin, et al.
Publicado: (2025)
por: Hu, Jianmin, et al.
Publicado: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
por: Chen, Xing, et al.
Publicado: (2025)
por: Chen, Xing, et al.
Publicado: (2025)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
por: Wu, Feiyang, et al.
Publicado: (2025)
por: Wu, Feiyang, et al.
Publicado: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
por: Wilkins, Grant, et al.
Publicado: (2024)
por: Wilkins, Grant, et al.
Publicado: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
por: He, Wenhao, et al.
Publicado: (2026)
por: He, Wenhao, et al.
Publicado: (2026)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
por: Nguyen, Thanh-Tung, et al.
Publicado: (2025)
por: Nguyen, Thanh-Tung, et al.
Publicado: (2025)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)
por: Li, Haley, et al.
Publicado: (2026)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
por: Lin, Yanying, et al.
Publicado: (2025)
por: Lin, Yanying, et al.
Publicado: (2025)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
por: Kong, Z. Jonny, et al.
Publicado: (2025)
por: Kong, Z. Jonny, et al.
Publicado: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
por: Bai, Fan, et al.
Publicado: (2026)
por: Bai, Fan, et al.
Publicado: (2026)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
por: O'Quinn, Austin, et al.
Publicado: (2025)
por: O'Quinn, Austin, et al.
Publicado: (2025)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
por: Wu, Z., et al.
Publicado: (2025)
por: Wu, Z., et al.
Publicado: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
por: Liao, Junhan, et al.
Publicado: (2025)
por: Liao, Junhan, et al.
Publicado: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
por: Nie, Chengyi, et al.
Publicado: (2024)
por: Nie, Chengyi, et al.
Publicado: (2024)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
por: Wolfrath, Joel, et al.
Publicado: (2025)
por: Wolfrath, Joel, et al.
Publicado: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
por: Zhao, Zhixin, et al.
Publicado: (2026)
por: Zhao, Zhixin, et al.
Publicado: (2026)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022)
por: Yang, Xiang, et al.
Publicado: (2022)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
por: Sheng, Jinhao, et al.
Publicado: (2025)
por: Sheng, Jinhao, et al.
Publicado: (2025)
Ejemplares similares
-
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
por: Ahmad, Sohaib, et al.
Publicado: (2024) -
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
por: Yang, Qizheng, et al.
Publicado: (2025) -
TridentServe: A Stage-level Serving System for Diffusion Pipelines
por: Xia, Yifei, et al.
Publicado: (2025) -
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
por: Razavi, Kamran, et al.
Publicado: (2024) -
Declarative Data Pipeline for Large Scale ML Services
por: Yang, Yunzhao, et al.
Publicado: (2025)