Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Bin, Zhang, Chen, Peng, Tao, Zhao, Hanyu, Xiao, Wencong, Sun, Minmin, Liu, Anmin, Zhang, Zhipeng, Li, Lanbo, Qiu, Xiafei, Li, Shen, Ji, Zhigang, Xie, Tao, Li, Yong, Lin, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914858771939328
author Lin, Bin
Zhang, Chen
Peng, Tao
Zhao, Hanyu
Xiao, Wencong
Sun, Minmin
Liu, Anmin
Zhang, Zhipeng
Li, Lanbo
Qiu, Xiafei
Li, Shen
Ji, Zhigang
Xie, Tao
Li, Yong
Lin, Wei
author_facet Lin, Bin
Zhang, Chen
Peng, Tao
Zhao, Hanyu
Xiao, Wencong
Sun, Minmin
Liu, Anmin
Zhang, Zhipeng
Li, Lanbo
Qiu, Xiafei
Li, Shen
Ji, Zhigang
Xie, Tao
Li, Yong
Lin, Wei
contents Large Language Models (LLMs) demonstrate substantial potential across a diverse array of domains via request serving. However, as trends continue to push for expanding context sizes, the autoregressive nature of LLMs results in highly dynamic behavior of the attention layers, showcasing significant differences in computational characteristics and memory requirements from the non-attention layers. This presents substantial challenges for resource management and performance optimization in service systems. Existing static model parallelism and resource allocation strategies fall short when dealing with this dynamicity. To address the issue, we propose Infinite-LLM, a novel LLM serving system designed to effectively handle dynamic context lengths. Infinite-LLM disaggregates attention layers from an LLM's inference process, facilitating flexible and independent resource scheduling that optimizes computational performance and enhances memory utilization jointly. By leveraging a pooled GPU memory strategy across a cluster, Infinite-LLM not only significantly boosts system throughput but also supports extensive context lengths. Evaluated on a dataset with context lengths ranging from a few to 2000K tokens across a cluster with 32 A100 GPUs, Infinite-LLM demonstrates throughput improvement of 1.35-3.4x compared to state-of-the-art methods, enabling efficient and elastic LLM deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02669
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Lin, Bin
Zhang, Chen
Peng, Tao
Zhao, Hanyu
Xiao, Wencong
Sun, Minmin
Liu, Anmin
Zhang, Zhipeng
Li, Lanbo
Qiu, Xiafei
Li, Shen
Ji, Zhigang
Xie, Tao
Li, Yong
Lin, Wei
Distributed, Parallel, and Cluster Computing
Hardware Architecture
Large Language Models (LLMs) demonstrate substantial potential across a diverse array of domains via request serving. However, as trends continue to push for expanding context sizes, the autoregressive nature of LLMs results in highly dynamic behavior of the attention layers, showcasing significant differences in computational characteristics and memory requirements from the non-attention layers. This presents substantial challenges for resource management and performance optimization in service systems. Existing static model parallelism and resource allocation strategies fall short when dealing with this dynamicity. To address the issue, we propose Infinite-LLM, a novel LLM serving system designed to effectively handle dynamic context lengths. Infinite-LLM disaggregates attention layers from an LLM's inference process, facilitating flexible and independent resource scheduling that optimizes computational performance and enhances memory utilization jointly. By leveraging a pooled GPU memory strategy across a cluster, Infinite-LLM not only significantly boosts system throughput but also supports extensive context lengths. Evaluated on a dataset with context lengths ranging from a few to 2000K tokens across a cluster with 32 A100 GPUs, Infinite-LLM demonstrates throughput improvement of 1.35-3.4x compared to state-of-the-art methods, enabling efficient and elastic LLM deployment.
title Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
topic Distributed, Parallel, and Cluster Computing
Hardware Architecture
url https://arxiv.org/abs/2401.02669