ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Seungbeom, Goo, Jeonghoe, Jeon, Eunjoo, Yang, Mingyu, Jang, Minsung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910943431098368
author Choi, Seungbeom
Goo, Jeonghoe
Jeon, Eunjoo
Yang, Mingyu
Jang, Minsung
author_facet Choi, Seungbeom
Goo, Jeonghoe
Jeon, Eunjoo
Yang, Mingyu
Jang, Minsung
contents We propose ELIS, a serving system for Large Language Models (LLMs) featuring an Iterative Shortest Remaining Time First (ISRTF) scheduler designed to efficiently manage inference tasks with the shortest remaining tokens. Current LLM serving systems often employ a first-come-first-served scheduling strategy, which can lead to the "head-of-line blocking" problem. To overcome this limitation, it is necessary to predict LLM inference times and apply a shortest job first scheduling strategy. However, due to the auto-regressive nature of LLMs, predicting the inference latency is challenging. ELIS addresses this challenge by training a response length predictor for LLMs using the BGE model, an encoder-based state-of-the-art model. Additionally, we have devised the ISRTF scheduling strategy, an optimization of shortest remaining time first tailored to existing LLM iteration batching. To evaluate our work in an industrial setting, we simulate streams of requests based on our study of real-world user LLM serving trace records. Furthermore, we implemented ELIS as a cloud-native scheduler system on Kubernetes to evaluate its performance in production environments. Our experimental results demonstrate that ISRTF reduces the average job completion time by up to 19.6%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
Choi, Seungbeom
Goo, Jeonghoe
Jeon, Eunjoo
Yang, Mingyu
Jang, Minsung
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
We propose ELIS, a serving system for Large Language Models (LLMs) featuring an Iterative Shortest Remaining Time First (ISRTF) scheduler designed to efficiently manage inference tasks with the shortest remaining tokens. Current LLM serving systems often employ a first-come-first-served scheduling strategy, which can lead to the "head-of-line blocking" problem. To overcome this limitation, it is necessary to predict LLM inference times and apply a shortest job first scheduling strategy. However, due to the auto-regressive nature of LLMs, predicting the inference latency is challenging. ELIS addresses this challenge by training a response length predictor for LLMs using the BGE model, an encoder-based state-of-the-art model. Additionally, we have devised the ISRTF scheduling strategy, an optimization of shortest remaining time first tailored to existing LLM iteration batching. To evaluate our work in an industrial setting, we simulate streams of requests based on our study of real-world user LLM serving trace records. Furthermore, we implemented ELIS as a cloud-native scheduler system on Kubernetes to evaluate its performance in production environments. Our experimental results demonstrate that ISRTF reduces the average job completion time by up to 19.6%.
title ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.09142