Saved in:
Bibliographic Details
Main Authors: Pan, Xiurui, Li, Endian, Li, Qiao, Liang, Shengwen, Shan, Yizhou, Zhou, Ke, Luo, Yingwei, Wang, Xiaolin, Zhang, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.04992
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910594814181376
author Pan, Xiurui
Li, Endian
Li, Qiao
Liang, Shengwen
Shan, Yizhou
Zhou, Ke
Luo, Yingwei
Wang, Xiaolin
Zhang, Jie
author_facet Pan, Xiurui
Li, Endian
Li, Qiao
Liang, Shengwen
Shan, Yizhou
Zhou, Ke
Luo, Yingwei
Wang, Xiaolin
Zhang, Jie
contents The widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constraint scenarios (e.g., edge computing and personal devices). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstInfer, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstInfer designs a dedicated flash-aware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstInfer improves throughput for long-sequence inference by up to 11.1$\times$, compared to existing SSD-based solutions such as FlexGen.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04992
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
Pan, Xiurui
Li, Endian
Li, Qiao
Liang, Shengwen
Shan, Yizhou
Zhou, Ke
Luo, Yingwei
Wang, Xiaolin
Zhang, Jie
Hardware Architecture
Computation and Language
The widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constraint scenarios (e.g., edge computing and personal devices). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstInfer, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstInfer designs a dedicated flash-aware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstInfer improves throughput for long-sequence inference by up to 11.1$\times$, compared to existing SSD-based solutions such as FlexGen.
title InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
topic Hardware Architecture
Computation and Language
url https://arxiv.org/abs/2409.04992