NimbusGuard: A Novel Framework for Proactive Kubernetes Autoscaling Using Deep Q-Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wanigasooriya, Chamath, Ekanayake, Indrajith |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
por: Xie, Tianfang
Publicado: (2026)
por: Xie, Tianfang
Publicado: (2026)
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
por: Park, Gyeongseo, et al.
Publicado: (2026)
por: Park, Gyeongseo, et al.
Publicado: (2026)
Verifying In-Network Computing Systems for Design Risks
por: Bai, Tianyu, et al.
Publicado: (2026)
por: Bai, Tianyu, et al.
Publicado: (2026)
Using a Market Economy to Provision Compute Resources Across Planet-wide Clusters
por: Stokely, Murray, et al.
Publicado: (2025)
por: Stokely, Murray, et al.
Publicado: (2025)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
por: Li, Zhifei, et al.
Publicado: (2026)
por: Li, Zhifei, et al.
Publicado: (2026)
Shipwright: Proving liveness of distributed systems with Byzantine participants
por: Leung, Derek, et al.
Publicado: (2025)
por: Leung, Derek, et al.
Publicado: (2025)
OPTIMUMP2P: Fast and Reliable Gossiping in P2P Networks
por: Nicolaou, Nicolas, et al.
Publicado: (2025)
por: Nicolaou, Nicolas, et al.
Publicado: (2025)
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
por: Talluri, Sacheendra, et al.
Publicado: (2025)
por: Talluri, Sacheendra, et al.
Publicado: (2025)
Distributed Recoverable Sketches (Extended Version)
por: Cohen, Diana, et al.
Publicado: (2025)
por: Cohen, Diana, et al.
Publicado: (2025)
Intersections of Web3 and AI -- View in 2024
por: Hyland-Wood, David, et al.
Publicado: (2024)
por: Hyland-Wood, David, et al.
Publicado: (2024)
Artifact Evaluation for Distributed Systems: Current Practices and Beyond
por: Sedghpour, Mohammad Reza Saleh, et al.
Publicado: (2024)
por: Sedghpour, Mohammad Reza Saleh, et al.
Publicado: (2024)
NotebookOS: A Replicated Notebook Platform for Interactive Training with On-Demand GPUs
por: Carver, Benjamin, et al.
Publicado: (2025)
por: Carver, Benjamin, et al.
Publicado: (2025)
Dodoor: Efficient Randomized Decentralized Scheduling with Load Caching for Heterogeneous Tasks and Clusters
por: Da, Wei, et al.
Publicado: (2025)
por: Da, Wei, et al.
Publicado: (2025)
Generic Multicast (Extended Version)
por: Bolina, José Augusto, et al.
Publicado: (2024)
por: Bolina, José Augusto, et al.
Publicado: (2024)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
por: Khan, Shamsher
Publicado: (2026)
por: Khan, Shamsher
Publicado: (2026)
dpBento: Benchmarking DPUs for Data Processing
por: Hu, Jiasheng, et al.
Publicado: (2025)
por: Hu, Jiasheng, et al.
Publicado: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
por: Will, Jonathan, et al.
Publicado: (2025)
por: Will, Jonathan, et al.
Publicado: (2025)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
por: De Sensi, Daniele, et al.
Publicado: (2025)
por: De Sensi, Daniele, et al.
Publicado: (2025)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
por: Sidik, Bronislav, et al.
Publicado: (2026)
por: Sidik, Bronislav, et al.
Publicado: (2026)
ACME: Adaptive Customization of Large Models via Distributed Systems
por: Dai, Ziming, et al.
Publicado: (2025)
por: Dai, Ziming, et al.
Publicado: (2025)
CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
por: Zhang, Guilin, et al.
Publicado: (2026)
por: Zhang, Guilin, et al.
Publicado: (2026)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
por: Sunkara, Krishna Chaitanya
Publicado: (2026)
por: Sunkara, Krishna Chaitanya
Publicado: (2026)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
por: Dettinger, Falk, et al.
Publicado: (2025)
por: Dettinger, Falk, et al.
Publicado: (2025)
Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads
por: Stokely, Murray, et al.
Publicado: (2025)
por: Stokely, Murray, et al.
Publicado: (2025)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
por: Khodabandehlou, Mohammadali, et al.
Publicado: (2026)
por: Khodabandehlou, Mohammadali, et al.
Publicado: (2026)
GPUnion: Autonomous GPU Sharing on Campus
por: Li, Yufang, et al.
Publicado: (2025)
por: Li, Yufang, et al.
Publicado: (2025)
Alea-BFT: Practical Asynchronous Byzantine Fault Tolerance
por: Antunes, Diogo S., et al.
Publicado: (2024)
por: Antunes, Diogo S., et al.
Publicado: (2024)
Laminar: A Probe-First Scheduling Paradigm with Deterministic Runtime Survival
por: Chu, Zhengyan
Publicado: (2026)
por: Chu, Zhengyan
Publicado: (2026)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
por: Pan, Ding, et al.
Publicado: (2026)
por: Pan, Ding, et al.
Publicado: (2026)
Knowledge Graphs-Driven Intelligence for Distributed Decision Systems
por: Napoli, Rosario, et al.
Publicado: (2026)
por: Napoli, Rosario, et al.
Publicado: (2026)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
por: Dong, Fan, et al.
Publicado: (2023)
por: Dong, Fan, et al.
Publicado: (2023)
Pioplat: A Scalable, Low-Cost Framework for Latency Reduction in Ethereum Blockchain
por: Wang, Ke, et al.
Publicado: (2024)
por: Wang, Ke, et al.
Publicado: (2024)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
por: Bhattacharya, Siddhartha, et al.
Publicado: (2024)
por: Bhattacharya, Siddhartha, et al.
Publicado: (2024)
A Preliminary Model of Coordination-free Consistency
por: Li, Shulu, et al.
Publicado: (2025)
por: Li, Shulu, et al.
Publicado: (2025)
Scalable Genomic Context Analysis with GCsnap2 on HPC Clusters
por: Krummenacher, Reto, et al.
Publicado: (2025)
por: Krummenacher, Reto, et al.
Publicado: (2025)
Commitment Against Front Running Attacks
por: Canidio, Andrea, et al.
Publicado: (2023)
por: Canidio, Andrea, et al.
Publicado: (2023)
Parallel/Distributed Tabu Search for Scheduling Microprocessor Tasks in Hybrid Flowshop
por: Janiak, Adam, et al.
Publicado: (2025)
por: Janiak, Adam, et al.
Publicado: (2025)
Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud Queries
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
por: Pathak, Prashant Kumar, et al.
Publicado: (2026)
Nezha: Deployable and High-Performance Consensus Using Synchronized Clocks
por: Geng, Jinkun, et al.
Publicado: (2022)
por: Geng, Jinkun, et al.
Publicado: (2022)
Privacy-Aware Split Inference with Speculative Decoding for Large Language Models over Wide-Area Networks
por: Cunningham, Michael
Publicado: (2026)
por: Cunningham, Michael
Publicado: (2026)
Ejemplares similares
-
Rank-Aware Resource Scheduling for Tightly-Coupled MPI Workloads on Kubernetes
por: Xie, Tianfang
Publicado: (2026) -
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
por: Park, Gyeongseo, et al.
Publicado: (2026) -
Verifying In-Network Computing Systems for Design Risks
por: Bai, Tianyu, et al.
Publicado: (2026) -
Using a Market Economy to Provision Compute Resources Across Planet-wide Clusters
por: Stokely, Murray, et al.
Publicado: (2025) -
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
por: Li, Zhifei, et al.
Publicado: (2026)