Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914858771939328 |
|---|---|
| author | Lin, Bin Zhang, Chen Peng, Tao Zhao, Hanyu Xiao, Wencong Sun, Minmin Liu, Anmin Zhang, Zhipeng Li, Lanbo Qiu, Xiafei Li, Shen Ji, Zhigang Xie, Tao Li, Yong Lin, Wei |
| author_facet | Lin, Bin Zhang, Chen Peng, Tao Zhao, Hanyu Xiao, Wencong Sun, Minmin Liu, Anmin Zhang, Zhipeng Li, Lanbo Qiu, Xiafei Li, Shen Ji, Zhigang Xie, Tao Li, Yong Lin, Wei |
| contents | Large Language Models (LLMs) demonstrate substantial potential across a diverse array of domains via request serving. However, as trends continue to push for expanding context sizes, the autoregressive nature of LLMs results in highly dynamic behavior of the attention layers, showcasing significant differences in computational characteristics and memory requirements from the non-attention layers. This presents substantial challenges for resource management and performance optimization in service systems. Existing static model parallelism and resource allocation strategies fall short when dealing with this dynamicity. To address the issue, we propose Infinite-LLM, a novel LLM serving system designed to effectively handle dynamic context lengths. Infinite-LLM disaggregates attention layers from an LLM's inference process, facilitating flexible and independent resource scheduling that optimizes computational performance and enhances memory utilization jointly. By leveraging a pooled GPU memory strategy across a cluster, Infinite-LLM not only significantly boosts system throughput but also supports extensive context lengths. Evaluated on a dataset with context lengths ranging from a few to 2000K tokens across a cluster with 32 A100 GPUs, Infinite-LLM demonstrates throughput improvement of 1.35-3.4x compared to state-of-the-art methods, enabling efficient and elastic LLM deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_02669 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache Lin, Bin Zhang, Chen Peng, Tao Zhao, Hanyu Xiao, Wencong Sun, Minmin Liu, Anmin Zhang, Zhipeng Li, Lanbo Qiu, Xiafei Li, Shen Ji, Zhigang Xie, Tao Li, Yong Lin, Wei Distributed, Parallel, and Cluster Computing Hardware Architecture Large Language Models (LLMs) demonstrate substantial potential across a diverse array of domains via request serving. However, as trends continue to push for expanding context sizes, the autoregressive nature of LLMs results in highly dynamic behavior of the attention layers, showcasing significant differences in computational characteristics and memory requirements from the non-attention layers. This presents substantial challenges for resource management and performance optimization in service systems. Existing static model parallelism and resource allocation strategies fall short when dealing with this dynamicity. To address the issue, we propose Infinite-LLM, a novel LLM serving system designed to effectively handle dynamic context lengths. Infinite-LLM disaggregates attention layers from an LLM's inference process, facilitating flexible and independent resource scheduling that optimizes computational performance and enhances memory utilization jointly. By leveraging a pooled GPU memory strategy across a cluster, Infinite-LLM not only significantly boosts system throughput but also supports extensive context lengths. Evaluated on a dataset with context lengths ranging from a few to 2000K tokens across a cluster with 32 A100 GPUs, Infinite-LLM demonstrates throughput improvement of 1.35-3.4x compared to state-of-the-art methods, enabling efficient and elastic LLM deployment. |
| title | Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache |
| topic | Distributed, Parallel, and Cluster Computing Hardware Architecture |
| url | https://arxiv.org/abs/2401.02669 |