STAR: Decode-Phase Rescheduling for LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhibin, Hong, Zetao, Li, Xue, Wang, Zibo, Li, Shipeng, Meng, Qingkai, Wang, Qing, Huan, Chengying, Gu, Rong, Zhong, Sheng, Tian, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915974694830080
author Wang, Zhibin
Hong, Zetao
Li, Xue
Wang, Zibo
Li, Shipeng
Meng, Qingkai
Wang, Qing
Huan, Chengying
Gu, Rong
Zhong, Sheng
Tian, Chen
author_facet Wang, Zhibin
Hong, Zetao
Li, Xue
Wang, Zibo
Li, Shipeng
Meng, Qingkai
Wang, Qing
Huan, Chengying
Gu, Rong
Zhong, Sheng
Tian, Chen
contents Large Language Model (LLM) inference has emerged as a fundamental paradigm, however, variations in output length cause severe workload imbalance in the decode phase, particularly for long-output reasoning tasks. Existing systems, such as PD disaggregation architectures, rely on static prefill-to-decode scheduling, which often results in SLO violations and OOM failures under evolving decode workloads. In this paper, we propose STAR, a decode rescheduling system powered by length prediction to anticipate future workloads. Our core contributions include: (1) A lightweight and continuous LLM-native prediction method that leverages LLM hidden state to model remaining generation length with high precision (reducing MAE by 49.42%) and low overhead (cutting predictor parameters by 93.28%); (2) A rescheduling solution in decode phase with a dynamic balancing mechanism that integrates current and predicted workloads, reducing P99 TPOT by 75.1% and achieving 2.63 times higher goodput.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13668
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAR: Decode-Phase Rescheduling for LLM Inference
Wang, Zhibin
Hong, Zetao
Li, Xue
Wang, Zibo
Li, Shipeng
Meng, Qingkai
Wang, Qing
Huan, Chengying
Gu, Rong
Zhong, Sheng
Tian, Chen
Distributed, Parallel, and Cluster Computing
Machine Learning
Large Language Model (LLM) inference has emerged as a fundamental paradigm, however, variations in output length cause severe workload imbalance in the decode phase, particularly for long-output reasoning tasks. Existing systems, such as PD disaggregation architectures, rely on static prefill-to-decode scheduling, which often results in SLO violations and OOM failures under evolving decode workloads. In this paper, we propose STAR, a decode rescheduling system powered by length prediction to anticipate future workloads. Our core contributions include: (1) A lightweight and continuous LLM-native prediction method that leverages LLM hidden state to model remaining generation length with high precision (reducing MAE by 49.42%) and low overhead (cutting predictor parameters by 93.28%); (2) A rescheduling solution in decode phase with a dynamic balancing mechanism that integrates current and predicted workloads, reducing P99 TPOT by 75.1% and achieving 2.63 times higher goodput.
title STAR: Decode-Phase Rescheduling for LLM Inference
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2510.13668