Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
Fuente:
arXiv
Saved in:
| Main Authors: | Trappen, Tim, Keßler, Robert, Pabel, Roland, Achter, Viktor, Wesner, Stefan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
by: De Sensi, Daniele, et al.
Published: (2025)
by: De Sensi, Daniele, et al.
Published: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024)
by: Mukherjee, Dipayan, et al.
Published: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Comprehensive Plugin-Based Monitoring of Nexflow Workflow Executions
by: Kharma, Sami, et al.
Published: (2026)
by: Kharma, Sami, et al.
Published: (2026)
GPU-centric Communication Schemes for HPC and ML Applications
by: Namashivayam, Naveen
Published: (2025)
by: Namashivayam, Naveen
Published: (2025)
Application-Driven Exascale: The JUPITER Benchmark Suite
by: Herten, Andreas, et al.
Published: (2024)
by: Herten, Andreas, et al.
Published: (2024)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025)
by: Tharwani, Jay, et al.
Published: (2025)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
Intent-driven scheduling of backup jobs
by: Dutta, Souvik, et al.
Published: (2024)
by: Dutta, Souvik, et al.
Published: (2024)
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
by: De Sensi, Daniele, et al.
Published: (2024)
by: De Sensi, Daniele, et al.
Published: (2024)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
by: Shakerdargah, Mohammadali, et al.
Published: (2024)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
by: Mao, Ying, et al.
Published: (2020)
by: Mao, Ying, et al.
Published: (2020)
Usability Evaluation of Cloud for HPC Applications
by: Sochat, Vanessa, et al.
Published: (2025)
by: Sochat, Vanessa, et al.
Published: (2025)
Speed, power and cost implications for GPU acceleration of Computational Fluid Dynamics on HPC systems
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
by: Cooper-Baldock, Zachary, et al.
Published: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
by: Karfakis, George, et al.
Published: (2025)
by: Karfakis, George, et al.
Published: (2025)
Extrae.jl: Julia bindings for the Extrae HPC Profiler
by: Sanchez-Ramirez, Sergio, et al.
Published: (2025)
by: Sanchez-Ramirez, Sergio, et al.
Published: (2025)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery
by: Allcock, William E., et al.
Published: (2025)
by: Allcock, William E., et al.
Published: (2025)
LLload: Simplifying Real-Time Job Monitoring for HPC Users
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
by: Miletto, Marcelo Cogo, et al.
Published: (2024)
by: Miletto, Marcelo Cogo, et al.
Published: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Characterising resource management performance in Kubernetes
by: Medel, Víctor, et al.
Published: (2024)
by: Medel, Víctor, et al.
Published: (2024)
Operational Strategies for Non-Disruptive Scheduling Transitions in Production HPC Systems
by: MacLachlan, Glen, et al.
Published: (2026)
by: MacLachlan, Glen, et al.
Published: (2026)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
by: Lin, Wei-Chen, et al.
Published: (2024)
by: Lin, Wei-Chen, et al.
Published: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
NimbusGuard: A Novel Framework for Proactive Kubernetes Autoscaling Using Deep Q-Networks
by: Wanigasooriya, Chamath, et al.
Published: (2026)
by: Wanigasooriya, Chamath, et al.
Published: (2026)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026)
by: Arif, Moiz, et al.
Published: (2026)
Exploiting Spot Instances for Time-Critical Cloud Workloads Using Optimal Randomized Strategies
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Opportunistic Scheduling for Optimal Spot Instance Savings in the Cloud
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Modeling and Characterizing Service Interference in Dynamic Infrastructures
by: Medel, VÍctor, et al.
Published: (2024)
by: Medel, VÍctor, et al.
Published: (2024)
Scalability Evaluation of HPC Multi-GPU Training for ECG-based LLMs
by: Mileski, Dimitar, et al.
Published: (2025)
by: Mileski, Dimitar, et al.
Published: (2025)
Serverless Cold Starts and Where to Find Them
by: Joosen, Artjom, et al.
Published: (2024)
by: Joosen, Artjom, et al.
Published: (2024)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)
by: Jallouli, Chaimae, et al.
Published: (2026)
Similar Items
-
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
by: De Sensi, Daniele, et al.
Published: (2025) -
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025) -
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024) -
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)