Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Srivatsa, Vikranth, He, Zijian, Abhyankar, Reyna, Li, Dongming, Zhang, Yiying
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913528646991872
author Srivatsa, Vikranth
He, Zijian
Abhyankar, Reyna
Li, Dongming
Zhang, Yiying
author_facet Srivatsa, Vikranth
He, Zijian
Abhyankar, Reyna
Li, Dongming
Zhang, Yiying
contents Prompts to large language models (LLMs) have evolved beyond simple user questions. For LLMs to solve complex problems, today's practices are to include domain-specific instructions, illustration of tool usages, and/or long context such as textbook chapters in prompts. As such, many parts of prompts are repetitive across requests. Recent works propose to cache and reuse KV state of prompts. However, they are all confined to a single-GPU optimization, while production LLM serving systems are distributed by nature. This paper proposes Preble, the first distributed LLM serving platform that targets and optimizes for prompt sharing. We designed a distributed scheduling system that co-optimizes KV state reuse and computation load-balancing with a new scheduling algorithm and a hierarchical scheduling mechanism. Our evaluation of Preble with real workloads and request arrival patterns on two open-source LLMs shows that Preble outperforms the SOTA serving systems by 1.5X to 14.5X on average latency and 2X to 10X on p99 latency.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00023
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Preble: Efficient Distributed Prompt Scheduling for LLM Serving
Srivatsa, Vikranth
He, Zijian
Abhyankar, Reyna
Li, Dongming
Zhang, Yiying
Distributed, Parallel, and Cluster Computing
Machine Learning
Prompts to large language models (LLMs) have evolved beyond simple user questions. For LLMs to solve complex problems, today's practices are to include domain-specific instructions, illustration of tool usages, and/or long context such as textbook chapters in prompts. As such, many parts of prompts are repetitive across requests. Recent works propose to cache and reuse KV state of prompts. However, they are all confined to a single-GPU optimization, while production LLM serving systems are distributed by nature. This paper proposes Preble, the first distributed LLM serving platform that targets and optimizes for prompt sharing. We designed a distributed scheduling system that co-optimizes KV state reuse and computation load-balancing with a new scheduling algorithm and a hierarchical scheduling mechanism. Our evaluation of Preble with real workloads and request arrival patterns on two open-source LLMs shows that Preble outperforms the SOTA serving systems by 1.5X to 14.5X on average latency and 2X to 10X on p99 latency.
title Preble: Efficient Distributed Prompt Scheduling for LLM Serving
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2407.00023