SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Ke, Wang, Zhi, Hu, Wen, Yang, Tiannuo, Li, Jianguo, Zhang, Sheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
di: Gao, Wei, et al.
Pubblicazione: (2026)
di: Gao, Wei, et al.
Pubblicazione: (2026)
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
di: Hu, Tiancheng, et al.
Pubblicazione: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
di: Chow, Will
Pubblicazione: (2025)
di: Chow, Will
Pubblicazione: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
di: Oliaro, Gabriele, et al.
Pubblicazione: (2024)
di: Oliaro, Gabriele, et al.
Pubblicazione: (2024)
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
di: Egersdoerfer, Chris, et al.
Pubblicazione: (2026)
di: Egersdoerfer, Chris, et al.
Pubblicazione: (2026)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
di: Sedlak, Boris, et al.
Pubblicazione: (2024)
di: Sedlak, Boris, et al.
Pubblicazione: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
di: Jeon, Beomyeol, et al.
Pubblicazione: (2024)
di: Jeon, Beomyeol, et al.
Pubblicazione: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
di: Gu, Jianfeng, et al.
Pubblicazione: (2025)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
di: Ren, Feng, et al.
Pubblicazione: (2026)
di: Ren, Feng, et al.
Pubblicazione: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
di: Li, Luchang, et al.
Pubblicazione: (2026)
di: Li, Luchang, et al.
Pubblicazione: (2026)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
di: Yu, Jiahuan, et al.
Pubblicazione: (2026)
di: Yu, Jiahuan, et al.
Pubblicazione: (2026)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
di: Xuan, Mo, et al.
Pubblicazione: (2025)
di: Xuan, Mo, et al.
Pubblicazione: (2025)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
di: Du, Kuntai, et al.
Pubblicazione: (2025)
di: Du, Kuntai, et al.
Pubblicazione: (2025)
EC2MoE: Adaptive End-Cloud Pipeline Collaboration Enabling Scalable Mixture-of-Experts Inference
di: Yang, Zheming, et al.
Pubblicazione: (2025)
di: Yang, Zheming, et al.
Pubblicazione: (2025)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
di: Gao, Lei, et al.
Pubblicazione: (2024)
di: Gao, Lei, et al.
Pubblicazione: (2024)
Communication-Efficient Collaborative LLM Inference over LEO Satellite Networks
di: Zhang, Songge, et al.
Pubblicazione: (2026)
di: Zhang, Songge, et al.
Pubblicazione: (2026)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
di: Xu, Ao, et al.
Pubblicazione: (2025)
di: Xu, Ao, et al.
Pubblicazione: (2025)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
di: Peng, Haosong, et al.
Pubblicazione: (2024)
di: Peng, Haosong, et al.
Pubblicazione: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
di: Sun, Desen, et al.
Pubblicazione: (2025)
di: Sun, Desen, et al.
Pubblicazione: (2025)
Distributed Inference Performance Optimization for LLMs on CPUs
di: He, Pujiang, et al.
Pubblicazione: (2024)
di: He, Pujiang, et al.
Pubblicazione: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
di: Li, Zikun, et al.
Pubblicazione: (2025)
di: Li, Zikun, et al.
Pubblicazione: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
di: Yang, Zheming, et al.
Pubblicazione: (2025)
di: Yang, Zheming, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
di: Zhang, Zuoning, et al.
Pubblicazione: (2024)
di: Zhang, Zuoning, et al.
Pubblicazione: (2024)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
di: Du, Yin, et al.
Pubblicazione: (2026)
di: Du, Yin, et al.
Pubblicazione: (2026)
Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
di: Chen, Xiaopei, et al.
Pubblicazione: (2025)
di: Chen, Xiaopei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024) -
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
di: Gao, Wei, et al.
Pubblicazione: (2026) -
Enabling Efficient Batch Serving for LMaaS via Generation Length Prediction
di: Cheng, Ke, et al.
Pubblicazione: (2024) -
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
di: Hu, Tiancheng, et al.
Pubblicazione: (2026) -
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)