AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Guilin, Vippagunta, Srinivas, Nandagopal, Raghavendra, Raman, Suchitra, Xu, Jeff, Pfeiffer, Marcus, Chatterjee, Shreeshankar, Tan, Ziqi, Guo, Wulan, Jiang, Hailong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
von: Zhang, Guilin, et al.
Veröffentlicht: (2026)
von: Zhang, Guilin, et al.
Veröffentlicht: (2026)
Saarthi: An End-to-End Intelligent Platform for Optimising Distributed Serverless Workloads
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2025)
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
von: Park, Gijun
Veröffentlicht: (2026)
von: Park, Gijun
Veröffentlicht: (2026)
ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration
von: Baghel, Himanshu Singh
Veröffentlicht: (2026)
von: Baghel, Himanshu Singh
Veröffentlicht: (2026)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Implementation of New Security Features in CMSWEB Kubernetes Cluster at CERN
von: Ali, Aamir, et al.
Veröffentlicht: (2024)
von: Ali, Aamir, et al.
Veröffentlicht: (2024)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
von: Zhu, Botao, et al.
Veröffentlicht: (2025)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2025)
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2025)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
Signalling Health for Improved Kubernetes Microservice Availability
von: Roberts, Jacob, et al.
Veröffentlicht: (2025)
von: Roberts, Jacob, et al.
Veröffentlicht: (2025)
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
von: Xie, Tianfang
Veröffentlicht: (2026)
von: Xie, Tianfang
Veröffentlicht: (2026)
Readout-Side Bypass for Residual Hybrid Quantum-Classical Models
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Are Unikernels Ready for Serverless on the Edge?
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
von: Moebius, Felix, et al.
Veröffentlicht: (2024)
Enhancing Kubernetes Resilience through Anomaly Detection and Prediction
von: Anemogiannis, V., et al.
Veröffentlicht: (2025)
von: Anemogiannis, V., et al.
Veröffentlicht: (2025)
A Contention-Free Model for Converged Kubernetes on HPC
von: Sochat, Vanessa, et al.
Veröffentlicht: (2024)
von: Sochat, Vanessa, et al.
Veröffentlicht: (2024)
Multi-Event Triggers for Serverless Computing
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
Raptor: Distributed Scheduling for Serverless Functions
von: Exton, Kevin, et al.
Veröffentlicht: (2024)
von: Exton, Kevin, et al.
Veröffentlicht: (2024)
Energy Efficient Scheduling for Serverless Systems
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
Affinity-aware Serverless Function Scheduling
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
Litmus: Fair Pricing for Serverless Computing
von: Pei, Qi, et al.
Veröffentlicht: (2024)
von: Pei, Qi, et al.
Veröffentlicht: (2024)
Serverless Computing: Architecture, Concepts, and Applications
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
von: Weitzel, Derek, et al.
Veröffentlicht: (2025)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
von: Aslan, Haci Ismail, et al.
Veröffentlicht: (2025)
von: Aslan, Haci Ismail, et al.
Veröffentlicht: (2025)
A GPU-accelerated Molecular Docking Workflow with Kubernetes and Apache Airflow
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
von: Medeiros, Daniel, et al.
Veröffentlicht: (2024)
A User-centric Kubernetes-based Architecture for Green Cloud Computing
von: Zanotto, Matteo, et al.
Veröffentlicht: (2025)
von: Zanotto, Matteo, et al.
Veröffentlicht: (2025)
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
von: Chen, Zihao, et al.
Veröffentlicht: (2025)
FaasMeter: Energy-First Serverless Computing
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
Serverless Abstractions for Short-Running, Lightweight Streams
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
von: Carl, Natalie, et al.
Veröffentlicht: (2026)
Orchestrating the Execution of Serverless Functions in Hybrid Clouds
von: Peri, Aristotelis, et al.
Veröffentlicht: (2024)
von: Peri, Aristotelis, et al.
Veröffentlicht: (2024)
Konflux: Optimized Function Fusion for Serverless Applications
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
von: Kowallik, Niklas, et al.
Veröffentlicht: (2026)
Caching Aided Multi-Tenant Serverless Computing
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
Software Resource Disaggregation for HPC with Serverless Computing
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
von: Copik, Marcin, et al.
Veröffentlicht: (2024)
On the Complexity of Reachability Properties in Serverless Function Scheduling
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
von: Zhang, Guilin, et al.
Veröffentlicht: (2025) -
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025) -
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025) -
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
von: Zhang, Guilin, et al.
Veröffentlicht: (2025) -
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)