An SLO Driven and Cost-Aware Autoscaling Framework for Kubernetes
Fuente:
arXiv
Saved in:
| Main Authors: | Punniyamoorthy, Vinoth, Kumar, Bikesh, Saha, Sumit, Butra, Lokesh, Palanigounder, Mayilsamy, Agarwal, Akash Kumar, Kannan, Kabilan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
Push Down Optimization for Distributed Multi Cloud Data Integration
by: Kodali, Ravi Kiran, et al.
Published: (2026)
by: Kodali, Ravi Kiran, et al.
Published: (2026)
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
by: Tymoshenko, Ivan, et al.
Published: (2026)
by: Tymoshenko, Ivan, et al.
Published: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
by: Qi, S., et al.
Published: (2024)
by: Qi, S., et al.
Published: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024)
by: Hu, Kan, et al.
Published: (2024)
Container-level Energy Observability in Kubernetes Clusters
by: Pijnacker, Bjorn, et al.
Published: (2025)
by: Pijnacker, Bjorn, et al.
Published: (2025)
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
by: Shaikh, Faraz, et al.
Published: (2026)
by: Shaikh, Faraz, et al.
Published: (2026)
Scalable Cloud-Native Architectures for Intelligent PMU Data Processing
by: Chockalingam, Nachiappan, et al.
Published: (2025)
by: Chockalingam, Nachiappan, et al.
Published: (2025)
Optimizing OpenFaaS on Kubernetes: Comparative Analysis of Language Runtimes and Cluster Distributions
by: Ataie, Ehsan, et al.
Published: (2026)
by: Ataie, Ehsan, et al.
Published: (2026)
Simplifying Root Cause Analysis in Kubernetes with StateGraph and LLM
by: Xiang, Yong, et al.
Published: (2025)
by: Xiang, Yong, et al.
Published: (2025)
Building Castles in the Cloud: Architecting Resilient and Scalable Infrastructure
by: Gundla, Naresh Kumar
Published: (2024)
by: Gundla, Naresh Kumar
Published: (2024)
AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices
by: Ndadji, Brice Arléon Zemtsop, et al.
Published: (2025)
by: Ndadji, Brice Arléon Zemtsop, et al.
Published: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
SpotKube: Cost-Optimal Microservices Deployment with Cluster Autoscaling and Spot Pricing
by: Edirisinghe, Dasith, et al.
Published: (2024)
by: Edirisinghe, Dasith, et al.
Published: (2024)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
by: Singh, Samaresh Kumar, et al.
Published: (2026)
by: Singh, Samaresh Kumar, et al.
Published: (2026)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
AnTi-MiCS: Analytical Framework for Bounding Time in Embedded Mixed-Criticality Systems
by: Ranjbar, Behnaz, et al.
Published: (2026)
by: Ranjbar, Behnaz, et al.
Published: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
Cost-Effective Big Data Orchestration Using Dagster: A Multi-Platform Approach
by: Picatto, Hernan, et al.
Published: (2024)
by: Picatto, Hernan, et al.
Published: (2024)
Kubernetes in the Cloud vs. Bare Metal: A Comparative Study of Network Costs
by: Redoli, Rodrigo Mompo, et al.
Published: (2025)
by: Redoli, Rodrigo Mompo, et al.
Published: (2025)
Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure
by: Pagidoju, Ravi Teja
Published: (2026)
by: Pagidoju, Ravi Teja
Published: (2026)
A Privacy-Preserving Cloud Architecture for Distributed Machine Learning at Scale
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
by: Punniyamoorthy, Vinoth, et al.
Published: (2025)
Self-adaptive, Requirements-driven Autoscaling of Microservices
by: Nunes, João Paulo Karol Santos, et al.
Published: (2024)
by: Nunes, João Paulo Karol Santos, et al.
Published: (2024)
Proactive and Reactive Autoscaling Techniques for Edge Computing
by: Gupta, Suhrid, et al.
Published: (2025)
by: Gupta, Suhrid, et al.
Published: (2025)
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
by: Ringoot, Evelyne, et al.
Published: (2025)
by: Ringoot, Evelyne, et al.
Published: (2025)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
by: Aslan, Haci Ismail, et al.
Published: (2025)
by: Aslan, Haci Ismail, et al.
Published: (2025)
A Framework for Effective Invocation Methods of Various LLM Services
by: Wang, Can, et al.
Published: (2024)
by: Wang, Can, et al.
Published: (2024)
A Comprehensive Benchmarking Analysis of Fault Recovery in Stream Processing Frameworks
by: Vogel, Adriano, et al.
Published: (2024)
by: Vogel, Adriano, et al.
Published: (2024)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
by: Domke, Jens, et al.
Published: (2025)
by: Domke, Jens, et al.
Published: (2025)
FMI Meets SystemC: A Framework for Cross-Tool Virtual Prototyping
by: Bosbach, Nils, et al.
Published: (2025)
by: Bosbach, Nils, et al.
Published: (2025)
Implementing Multi-GPU Scientific Computing Miniapps Across Performance Portable Frameworks
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
A Comprehensive Experimentation Framework for Energy-Efficient Design of Cloud-Native Applications
by: Werner, Sebastian, et al.
Published: (2025)
by: Werner, Sebastian, et al.
Published: (2025)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
by: Kim, Taeyoon, et al.
Published: (2026)
by: Kim, Taeyoon, et al.
Published: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
by: Sun, Desen, et al.
Published: (2025)
by: Sun, Desen, et al.
Published: (2025)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
by: Seo, Eunil, et al.
Published: (2025)
by: Seo, Eunil, et al.
Published: (2025)
AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Similar Items
-
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
by: Punniyamoorthy, Vinoth, et al.
Published: (2025) -
Push Down Optimization for Distributed Multi Cloud Data Integration
by: Kodali, Ravi Kiran, et al.
Published: (2026) -
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity
by: Tymoshenko, Ivan, et al.
Published: (2026) -
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
by: Qi, S., et al.
Published: (2024) -
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
by: Hu, Kan, et al.
Published: (2024)