Teola: Towards End-to-End Optimization of LLM-based Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Xin, Jiang, Yimin, Yang, Yitao, Xu, Hong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915811403235328
author Tan, Xin
Jiang, Yimin
Yang, Yitao
Xu, Hong
author_facet Tan, Xin
Jiang, Yimin
Yang, Yitao
Xu, Hong
contents Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inference, end-to-end workflow optimization has been overlooked. Existing frameworks employ coarse-grained orchestration with task modules, which confines optimizations to within each module and yields suboptimal scheduling decisions. We propose fine-grained end-to-end orchestration, which utilizes task primitives as the basic units and represents each query's workflow as a primitive-level dataflow graph. This explicitly exposes a much larger design space, enables optimizations in parallelization and pipelining across primitives of different modules, and enhances scheduling to improve application-level performance. We build Teola, a novel orchestration framework for LLM-based applications that implements this scheme. Comprehensive experiments show that Teola can achieve up to 2.09x speedup over existing systems across various popular LLM applications. The code is available at https://github.com/NetX-lab/Ayo.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00326
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Teola: Towards End-to-End Optimization of LLM-based Applications
Tan, Xin
Jiang, Yimin
Yang, Yitao
Xu, Hong
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inference, end-to-end workflow optimization has been overlooked. Existing frameworks employ coarse-grained orchestration with task modules, which confines optimizations to within each module and yields suboptimal scheduling decisions. We propose fine-grained end-to-end orchestration, which utilizes task primitives as the basic units and represents each query's workflow as a primitive-level dataflow graph. This explicitly exposes a much larger design space, enables optimizations in parallelization and pipelining across primitives of different modules, and enhances scheduling to improve application-level performance. We build Teola, a novel orchestration framework for LLM-based applications that implements this scheme. Comprehensive experiments show that Teola can achieve up to 2.09x speedup over existing systems across various popular LLM applications. The code is available at https://github.com/NetX-lab/Ayo.
title Teola: Towards End-to-End Optimization of LLM-based Applications
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
url https://arxiv.org/abs/2407.00326