Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jain, Kunal, Parayil, Anjaly, Mallick, Ankur, Choukse, Esha, Qin, Xiaoting, Zhang, Jue, Goiri, Íñigo, Wang, Rujia, Bansal, Chetan, Rühle, Victor, Kulkarni, Anoop, Kofsky, Steve, Rajmohan, Saravan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917885976248320
author Jain, Kunal
Parayil, Anjaly
Mallick, Ankur
Choukse, Esha
Qin, Xiaoting
Zhang, Jue
Goiri, Íñigo
Wang, Rujia
Bansal, Chetan
Rühle, Victor
Kulkarni, Anoop
Kofsky, Steve
Rajmohan, Saravan
author_facet Jain, Kunal
Parayil, Anjaly
Mallick, Ankur
Choukse, Esha
Qin, Xiaoting
Zhang, Jue
Goiri, Íñigo
Wang, Rujia
Bansal, Chetan
Rühle, Victor
Kulkarni, Anoop
Kofsky, Steve
Rajmohan, Saravan
contents Large Language Model (LLM) workloads have distinct prefill and decode phases with different compute and memory requirements which should ideally be accounted for when scheduling input queries across different LLM instances in a cluster. However existing scheduling algorithms treat LLM workloads as monolithic jobs without considering the distinct characteristics of the two phases in each workload. This leads to sub-optimal scheduling and increased response latency. In this work, we start by characterizing factors affecting the response latency during LLM inference serving. We establish that better load balancing of inference requests across the available LLM instances can improve the end-to-end latency to a larger extent than merely focusing on optimizing the instance-level scheduler. Motivated by our findings, we propose a heuristic-guided reinforcement learning-based intelligent router for data-driven and workload-aware scheduling. Our router schedules queries across LLM instances by leveraging a trainable response-length predictor, and a novel formulation for estimating the impact of mixing different workloads and achieves over 11% lower end-to-end latency than existing approaches on a mix of public datasets and 7.8% lower end-to-end latency on real workload data with diverse input and output trends from Cloud Provider X. Additionally, the proposed framework can also serve as a standard for benchmarking different LLM inference schedulers since it provides the best latency for a given model, hardware, and instance-level scheduler combination.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13510
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
Jain, Kunal
Parayil, Anjaly
Mallick, Ankur
Choukse, Esha
Qin, Xiaoting
Zhang, Jue
Goiri, Íñigo
Wang, Rujia
Bansal, Chetan
Rühle, Victor
Kulkarni, Anoop
Kofsky, Steve
Rajmohan, Saravan
Distributed, Parallel, and Cluster Computing
Systems and Control
Large Language Model (LLM) workloads have distinct prefill and decode phases with different compute and memory requirements which should ideally be accounted for when scheduling input queries across different LLM instances in a cluster. However existing scheduling algorithms treat LLM workloads as monolithic jobs without considering the distinct characteristics of the two phases in each workload. This leads to sub-optimal scheduling and increased response latency. In this work, we start by characterizing factors affecting the response latency during LLM inference serving. We establish that better load balancing of inference requests across the available LLM instances can improve the end-to-end latency to a larger extent than merely focusing on optimizing the instance-level scheduler. Motivated by our findings, we propose a heuristic-guided reinforcement learning-based intelligent router for data-driven and workload-aware scheduling. Our router schedules queries across LLM instances by leveraging a trainable response-length predictor, and a novel formulation for estimating the impact of mixing different workloads and achieves over 11% lower end-to-end latency than existing approaches on a mix of public datasets and 7.8% lower end-to-end latency on real workload data with diverse input and output trends from Cloud Provider X. Additionally, the proposed framework can also serve as a standard for benchmarking different LLM inference schedulers since it provides the best latency for a given model, hardware, and instance-level scheduler combination.
title Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
topic Distributed, Parallel, and Cluster Computing
Systems and Control
url https://arxiv.org/abs/2408.13510