Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yechen, Kong, Xinhao, Chen, Tingjun, Zhuo, Danyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913377568161792
author Xu, Yechen
Kong, Xinhao
Chen, Tingjun
Zhuo, Danyang
author_facet Xu, Yechen
Kong, Xinhao
Chen, Tingjun
Zhuo, Danyang
contents The complexity of large language model (LLM) serving workloads has substantially increased due to the integration with external tool invocations, such as ChatGPT plugins. In this paper, we identify a new opportunity for efficient LLM serving for requests that trigger tools: tool partial execution alongside LLM decoding. To this end, we design Conveyor, an efficient LLM serving system optimized for handling requests involving external tools. We introduce a novel interface for tool developers to expose partial execution opportunities to the LLM serving system and a request scheduler that facilitates partial tool execution. Our results demonstrate that tool partial execution can improve request completion latency by up to 38.8%.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00059
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
Xu, Yechen
Kong, Xinhao
Chen, Tingjun
Zhuo, Danyang
Computation and Language
Distributed, Parallel, and Cluster Computing
Machine Learning
The complexity of large language model (LLM) serving workloads has substantially increased due to the integration with external tool invocations, such as ChatGPT plugins. In this paper, we identify a new opportunity for efficient LLM serving for requests that trigger tools: tool partial execution alongside LLM decoding. To this end, we design Conveyor, an efficient LLM serving system optimized for handling requests involving external tools. We introduce a novel interface for tool developers to expose partial execution opportunities to the LLM serving system and a request scheduler that facilitates partial tool execution. Our results demonstrate that tool partial execution can improve request completion latency by up to 38.8%.
title Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
topic Computation and Language
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2406.00059