Saved in:
Bibliographic Details
Main Authors: Lin, Chaofan, Han, Zhenhua, Zhang, Chengruidong, Yang, Yuqing, Yang, Fan, Chen, Chen, Qiu, Lili
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.19888
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909213307961344
author Lin, Chaofan
Han, Zhenhua
Zhang, Chengruidong
Yang, Yuqing
Yang, Fan
Chen, Chen
Qiu, Lili
author_facet Lin, Chaofan
Han, Zhenhua
Zhang, Chengruidong
Yang, Yuqing
Yang, Fan
Chen, Chen
Qiu, Lili
contents The rise of large language models (LLMs) has enabled LLM-based applications (a.k.a. AI agents or co-pilots), a new software paradigm that combines the strength of LLM and conventional software. Diverse LLM applications from different tenants could design complex workflows using multiple LLM requests to accomplish one task. However, they have to use the over-simplified request-level API provided by today's public LLM services, losing essential application-level information. Public LLM services have to blindly optimize individual LLM requests, leading to sub-optimal end-to-end performance of LLM applications. This paper introduces Parrot, an LLM service system that focuses on the end-to-end experience of LLM-based applications. Parrot proposes Semantic Variable, a unified abstraction to expose application-level knowledge to public LLM services. A Semantic Variable annotates an input/output variable in the prompt of a request, and creates the data pipeline when connecting multiple LLM requests, providing a natural way to program LLM applications. Exposing Semantic Variables to the public LLM service allows it to perform conventional data flow analysis to uncover the correlation across multiple LLM requests. This correlation opens a brand-new optimization space for the end-to-end performance of LLM-based applications. Extensive evaluations demonstrate that Parrot can achieve up to an order-of-magnitude improvement for popular and practical use cases of LLM applications.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19888
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
Lin, Chaofan
Han, Zhenhua
Zhang, Chengruidong
Yang, Yuqing
Yang, Fan
Chen, Chen
Qiu, Lili
Machine Learning
Artificial Intelligence
The rise of large language models (LLMs) has enabled LLM-based applications (a.k.a. AI agents or co-pilots), a new software paradigm that combines the strength of LLM and conventional software. Diverse LLM applications from different tenants could design complex workflows using multiple LLM requests to accomplish one task. However, they have to use the over-simplified request-level API provided by today's public LLM services, losing essential application-level information. Public LLM services have to blindly optimize individual LLM requests, leading to sub-optimal end-to-end performance of LLM applications. This paper introduces Parrot, an LLM service system that focuses on the end-to-end experience of LLM-based applications. Parrot proposes Semantic Variable, a unified abstraction to expose application-level knowledge to public LLM services. A Semantic Variable annotates an input/output variable in the prompt of a request, and creates the data pipeline when connecting multiple LLM requests, providing a natural way to program LLM applications. Exposing Semantic Variables to the public LLM service allows it to perform conventional data flow analysis to uncover the correlation across multiple LLM requests. This correlation opens a brand-new optimization space for the end-to-end performance of LLM-based applications. Extensive evaluations demonstrate that Parrot can achieve up to an order-of-magnitude improvement for popular and practical use cases of LLM applications.
title Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.19888