Ensuring Fair LLM Serving Amid Diverse Applications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Khan, Redwan Ibne Seraj, Jain, Kunal, Shen, Haiying, Mallick, Ankur, Parayil, Anjaly, Kulkarni, Anoop, Kofsky, Steve, Choudhary, Pankhuri, Amant, Renèe St., Wang, Rujia, Cheng, Yue, Butt, Ali R., Rühle, Victor, Bansal, Chetan, Rajmohan, Saravan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916494169866240
author Khan, Redwan Ibne Seraj
Jain, Kunal
Shen, Haiying
Mallick, Ankur
Parayil, Anjaly
Kulkarni, Anoop
Kofsky, Steve
Choudhary, Pankhuri
Amant, Renèe St.
Wang, Rujia
Cheng, Yue
Butt, Ali R.
Rühle, Victor
Bansal, Chetan
Rajmohan, Saravan
author_facet Khan, Redwan Ibne Seraj
Jain, Kunal
Shen, Haiying
Mallick, Ankur
Parayil, Anjaly
Kulkarni, Anoop
Kofsky, Steve
Choudhary, Pankhuri
Amant, Renèe St.
Wang, Rujia
Cheng, Yue
Butt, Ali R.
Rühle, Victor
Bansal, Chetan
Rajmohan, Saravan
contents In a multi-tenant large language model (LLM) serving platform hosting diverse applications, some users may submit an excessive number of requests, causing the service to become unavailable to other users and creating unfairness. Existing fairness approaches do not account for variations in token lengths across applications and multiple LLM calls, making them unsuitable for such platforms. To address the fairness challenge, this paper analyzes millions of requests from thousands of users on MS CoPilot, a real-world multi-tenant LLM platform hosted by Microsoft. Our analysis confirms the inadequacy of existing methods and guides the development of FairServe, a system that ensures fair LLM access across diverse applications. FairServe proposes application-characteristic aware request throttling coupled with a weighted service counter based scheduling technique to curb abusive behavior and ensure fairness. Our experimental results on real-world traces demonstrate FairServe's superior performance compared to the state-of-the-art method in ensuring fairness. We are actively working on deploying our system in production, expecting to benefit millions of customers world-wide.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15997
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ensuring Fair LLM Serving Amid Diverse Applications
Khan, Redwan Ibne Seraj
Jain, Kunal
Shen, Haiying
Mallick, Ankur
Parayil, Anjaly
Kulkarni, Anoop
Kofsky, Steve
Choudhary, Pankhuri
Amant, Renèe St.
Wang, Rujia
Cheng, Yue
Butt, Ali R.
Rühle, Victor
Bansal, Chetan
Rajmohan, Saravan
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Multiagent Systems
In a multi-tenant large language model (LLM) serving platform hosting diverse applications, some users may submit an excessive number of requests, causing the service to become unavailable to other users and creating unfairness. Existing fairness approaches do not account for variations in token lengths across applications and multiple LLM calls, making them unsuitable for such platforms. To address the fairness challenge, this paper analyzes millions of requests from thousands of users on MS CoPilot, a real-world multi-tenant LLM platform hosted by Microsoft. Our analysis confirms the inadequacy of existing methods and guides the development of FairServe, a system that ensures fair LLM access across diverse applications. FairServe proposes application-characteristic aware request throttling coupled with a weighted service counter based scheduling technique to curb abusive behavior and ensure fairness. Our experimental results on real-world traces demonstrate FairServe's superior performance compared to the state-of-the-art method in ensuring fairness. We are actively working on deploying our system in production, expecting to benefit millions of customers world-wide.
title Ensuring Fair LLM Serving Amid Diverse Applications
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Multiagent Systems
url https://arxiv.org/abs/2411.15997