The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kondrashov, Leonid, Zhou, Boxi, Wang, Hancheng, Ustiugov, Dmitrii |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
von: Kondrashov, Leonid, et al.
Veröffentlicht: (2025)
von: Kondrashov, Leonid, et al.
Veröffentlicht: (2025)
Shattering the Ephemeral Storage Cost Barrier for Data-Intensive Serverless Workflows
von: Ustiugov, Dmitrii, et al.
Veröffentlicht: (2023)
von: Ustiugov, Dmitrii, et al.
Veröffentlicht: (2023)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
von: Pan, Ding, et al.
Veröffentlicht: (2026)
von: Pan, Ding, et al.
Veröffentlicht: (2026)
NBI-Slurm: Simplified submission of Slurm jobs with energy saving mode
von: Telatin, Andrea
Veröffentlicht: (2026)
von: Telatin, Andrea
Veröffentlicht: (2026)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
von: Tsai, Yu-Shiang, et al.
Veröffentlicht: (2024)
von: Tsai, Yu-Shiang, et al.
Veröffentlicht: (2024)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
von: Sada, Mohammad Firas, et al.
Veröffentlicht: (2025)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
von: Jallouli, Chaimae, et al.
Veröffentlicht: (2026)
Big Data Workload Profiling for Energy-Aware Cloud Resource Management
von: Parikh, Milan, et al.
Veröffentlicht: (2026)
von: Parikh, Milan, et al.
Veröffentlicht: (2026)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
von: Lysenstøen, Christian
Veröffentlicht: (2026)
von: Lysenstøen, Christian
Veröffentlicht: (2026)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
von: Hu, Zhibo, et al.
Veröffentlicht: (2024)
von: Hu, Zhibo, et al.
Veröffentlicht: (2024)
Work-Efficient Parallel Non-Maximum Suppression Kernels
von: Oro, David, et al.
Veröffentlicht: (2025)
von: Oro, David, et al.
Veröffentlicht: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
von: Putra, Jody Almaida
Veröffentlicht: (2026)
von: Putra, Jody Almaida
Veröffentlicht: (2026)
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Nexus: Transparent I/O Offloading for High-Density Serverless Computing
von: Park, JooYoung, et al.
Veröffentlicht: (2026)
von: Park, JooYoung, et al.
Veröffentlicht: (2026)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
von: Rao, Swanand, et al.
Veröffentlicht: (2026)
von: Rao, Swanand, et al.
Veröffentlicht: (2026)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
Serverless Cold Starts and Where to Find Them
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
von: Joosen, Artjom, et al.
Veröffentlicht: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
von: Khan, Shamsher
Veröffentlicht: (2026)
von: Khan, Shamsher
Veröffentlicht: (2026)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
Attack-Centric by Design: A Program-Structure Taxonomy of Smart Contract Vulnerabilities
von: Hedayatnia, Parsa, et al.
Veröffentlicht: (2025)
von: Hedayatnia, Parsa, et al.
Veröffentlicht: (2025)
Evaluating the Overhead of the Performance Profiler Cloudprofiler With MooBench
von: Yang, Shinhyung, et al.
Veröffentlicht: (2024)
von: Yang, Shinhyung, et al.
Veröffentlicht: (2024)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
von: Staylor, Mills, et al.
Veröffentlicht: (2025)
von: Staylor, Mills, et al.
Veröffentlicht: (2025)
SpotKube: Cost-Optimal Microservices Deployment with Cluster Autoscaling and Spot Pricing
von: Edirisinghe, Dasith, et al.
Veröffentlicht: (2024)
von: Edirisinghe, Dasith, et al.
Veröffentlicht: (2024)
Analysis of Design Patterns and Benchmark Practices in Apache Kafka Event-Streaming Systems
von: Mohammad, Muzeeb
Veröffentlicht: (2025)
von: Mohammad, Muzeeb
Veröffentlicht: (2025)
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
Planetary computing for data-driven environmental policy-making
von: Ferris, Patrick, et al.
Veröffentlicht: (2023)
von: Ferris, Patrick, et al.
Veröffentlicht: (2023)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
von: Punniyamoorthy, Vinoth, et al.
Veröffentlicht: (2025)
Semaphores Augmented with a Waiting Array
von: Dice, Dave, et al.
Veröffentlicht: (2025)
von: Dice, Dave, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Melding the Serverless Control Plane with the Conventional Cluster Manager for Speed and Resource Efficiency
von: Kondrashov, Leonid, et al.
Veröffentlicht: (2025) -
Shattering the Ephemeral Storage Cost Barrier for Data-Intensive Serverless Workflows
von: Ustiugov, Dmitrii, et al.
Veröffentlicht: (2023) -
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
von: Pan, Ding, et al.
Veröffentlicht: (2026) -
NBI-Slurm: Simplified submission of Slurm jobs with energy saving mode
von: Telatin, Andrea
Veröffentlicht: (2026) -
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
von: Tsai, Yu-Shiang, et al.
Veröffentlicht: (2024)