Towards Resource-Efficient Serverless LLM Inference with SLINFER

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Chuhao, Li, Zijun, Chen, Quan, Zhao, Han, Tang, Xueyan, Guo, Minyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917145984630784
author Xu, Chuhao
Li, Zijun
Chen, Quan
Zhao, Han
Tang, Xueyan
Guo, Minyi
author_facet Xu, Chuhao
Li, Zijun
Chen, Quan
Zhao, Han
Tang, Xueyan
Guo, Minyi
contents The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow exclusive GPU allocation, we take a step back to explore modern platforms and find that: Emerging CPU architectures with built-in accelerators are capable of serving LLMs but remain underutilized, and both CPUs and GPUs can accommodate multiple LLMs simultaneously. We propose SLINFER, a resource-efficient serverless inference scheme tailored for small- to mid-sized LLMs that enables elastic and on-demand sharing across heterogeneous hardware. SLINFER tackles three fundamental challenges: (1) precise, fine-grained compute resource allocation at token-level to handle fluctuating computational demands; (2) a coordinated and forward-looking memory scaling mechanism to detect out-of-memory hazards and reduce operational overhead; and (3) a dual approach that consolidates fragmented instances through proactive preemption and reactive bin-packing. Experimental results on 4 32-core CPUs and 4 A100 GPUs show that SLINFER improves serving capacity by 47% - 62% through sharing, while further leveraging CPUs boosts this to 86% - 154%.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Resource-Efficient Serverless LLM Inference with SLINFER
Xu, Chuhao
Li, Zijun
Chen, Quan
Zhao, Han
Tang, Xueyan
Guo, Minyi
Distributed, Parallel, and Cluster Computing
The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow exclusive GPU allocation, we take a step back to explore modern platforms and find that: Emerging CPU architectures with built-in accelerators are capable of serving LLMs but remain underutilized, and both CPUs and GPUs can accommodate multiple LLMs simultaneously. We propose SLINFER, a resource-efficient serverless inference scheme tailored for small- to mid-sized LLMs that enables elastic and on-demand sharing across heterogeneous hardware. SLINFER tackles three fundamental challenges: (1) precise, fine-grained compute resource allocation at token-level to handle fluctuating computational demands; (2) a coordinated and forward-looking memory scaling mechanism to detect out-of-memory hazards and reduce operational overhead; and (3) a dual approach that consolidates fragmented instances through proactive preemption and reactive bin-packing. Experimental results on 4 32-core CPUs and 4 A100 GPUs show that SLINFER improves serving capacity by 47% - 62% through sharing, while further leveraging CPUs boosts this to 86% - 154%.
title Towards Resource-Efficient Serverless LLM Inference with SLINFER
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2507.00507