Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trappen, Tim, Keßler, Robert, Pabel, Roland, Achter, Viktor, Wesner, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912730640809984
author Trappen, Tim
Keßler, Robert
Pabel, Roland
Achter, Viktor
Wesner, Stefan
author_facet Trappen, Tim
Keßler, Robert
Pabel, Roland
Achter, Viktor
Wesner, Stefan
contents Due to rising demands for Artificial Inteligence (AI) inference, especially in higher education, novel solutions utilising existing infrastructure are emerging. The utilisation of High-Performance Computing (HPC) has become a prevalent approach for the implementation of such solutions. However, the classical operating model of HPC does not adapt well to the requirements of synchronous, user-facing dynamic AI application workloads. In this paper, we propose our solution that serves LLMs by integrating vLLM, Slurm and Kubernetes on the supercomputer \textit{RAMSES}. The initial benchmark indicates that the proposed architecture scales efficiently for 100, 500 and 1000 concurrent requests, incurring only an overhead of approximately 500 ms in terms of end-to-end latency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
Trappen, Tim
Keßler, Robert
Pabel, Roland
Achter, Viktor
Wesner, Stefan
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Databases
Performance
C.3; C.4; C.5.1
Due to rising demands for Artificial Inteligence (AI) inference, especially in higher education, novel solutions utilising existing infrastructure are emerging. The utilisation of High-Performance Computing (HPC) has become a prevalent approach for the implementation of such solutions. However, the classical operating model of HPC does not adapt well to the requirements of synchronous, user-facing dynamic AI application workloads. In this paper, we propose our solution that serves LLMs by integrating vLLM, Slurm and Kubernetes on the supercomputer \textit{RAMSES}. The initial benchmark indicates that the proposed architecture scales efficiently for 100, 500 and 1000 concurrent requests, incurring only an overhead of approximately 500 ms in terms of end-to-end latency.
title Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Databases
Performance
C.3; C.4; C.5.1
url https://arxiv.org/abs/2511.21413