SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
Fuente:
arXiv
Saved in:
| Main Author: | Lysenstøen, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)
by: Jallouli, Chaimae, et al.
Published: (2026)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
by: Arafat, Jahidul, et al.
Published: (2025)
by: Arafat, Jahidul, et al.
Published: (2025)
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026)
by: Šabanović, Ahmed, et al.
Published: (2026)
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)
by: Lubonja, Ariel, et al.
Published: (2026)
Intelligent Cloud Orchestration: A Hybrid Predictive and Heuristic Framework for Cost Optimization
by: Nagoriya, Heet, et al.
Published: (2026)
by: Nagoriya, Heet, et al.
Published: (2026)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
by: Gao, Yunquan, et al.
Published: (2025)
by: Gao, Yunquan, et al.
Published: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
Feature-Aware Task-to-Core Allocation in Embedded Multi-core Platforms via Statistical Learning
by: Pivezhandi, Mohammad, et al.
Published: (2025)
by: Pivezhandi, Mohammad, et al.
Published: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)
by: Pan, Ding, et al.
Published: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
by: Qi, S., et al.
Published: (2024)
by: Qi, S., et al.
Published: (2024)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
by: Yang, Haotian, et al.
Published: (2025)
by: Yang, Haotian, et al.
Published: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
TStore: Rethinking AI Model Hub with Tensor-Centric Compression
by: Lan, Tingfeng, et al.
Published: (2026)
by: Lan, Tingfeng, et al.
Published: (2026)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
by: Topcu, Burak, et al.
Published: (2026)
by: Topcu, Burak, et al.
Published: (2026)
CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
by: Zhang, Guilin, et al.
Published: (2026)
by: Zhang, Guilin, et al.
Published: (2026)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Heterogeneous Model Fusion for Privacy-Aware Multi-Camera Surveillance via Synthetic Domain Adaptation
by: Lu, Peggy Joy, et al.
Published: (2026)
by: Lu, Peggy Joy, et al.
Published: (2026)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
by: Gao, Yuxuan, et al.
Published: (2026)
by: Gao, Yuxuan, et al.
Published: (2026)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
by: Sada, Mohammad Firas, et al.
Published: (2025)
by: Sada, Mohammad Firas, et al.
Published: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025)
by: Khan, Kabir, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
HiDVFS: A Hierarchical Multi-Agent DVFS Scheduler for OpenMP DAG Workloads
by: Pivezhandi, Mohammad, et al.
Published: (2026)
by: Pivezhandi, Mohammad, et al.
Published: (2026)
LAMMPS-KOKKOS: Performance Portable Molecular Dynamics Across Exascale Architectures
by: Johansson, Anders, et al.
Published: (2025)
by: Johansson, Anders, et al.
Published: (2025)
Work-Efficient Parallel Non-Maximum Suppression Kernels
by: Oro, David, et al.
Published: (2025)
by: Oro, David, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
SLOs-Serve: Optimized Serving of Multi-SLO LLMs
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems
by: Mambelli, Marco, et al.
Published: (2025)
by: Mambelli, Marco, et al.
Published: (2025)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
How Can We Train Deep Learning Models Across Clouds and Continents? An Experimental Study
by: Erben, Alexander, et al.
Published: (2023)
by: Erben, Alexander, et al.
Published: (2023)
Similar Items
-
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026) -
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026) -
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
by: Arafat, Jahidul, et al.
Published: (2025) -
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026) -
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)