A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Hongsun, Song, Jaeyong, Shin, Changmin, Noh, Si Ung, Jung, Jaewon, Park, Jisung, Lee, Jinho
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912883349127168
author Jang, Hongsun
Song, Jaeyong
Shin, Changmin
Noh, Si Ung
Jung, Jaewon
Park, Jisung
Lee, Jinho
author_facet Jang, Hongsun
Song, Jaeyong
Shin, Changmin
Noh, Si Ung
Jung, Jaewon
Park, Jisung
Lee, Jinho
contents The computational and memory demands of large language models for generative inference present significant challenges for practical deployment. One promising solution targeting offline inference is offloading-based batched inference, which extends the GPU's memory hierarchy with host memory and storage. However, it often suffers from substantial I/O overhead, primarily due to the large KV cache sizes that scale with batch size and context window length. In this paper, we introduce HILOS, a framework that boosts offline inference throughput using near-storage processing. The core of HILOS is attention near storage, which offloads memory-intensive attention operations to near-storage accelerators, reducing traffic across the system interconnect. Building on attention near storage, HILOS incorporates three additional optimizations. First, cooperative X-cache minimizes KV cache I/O by exploiting available host resources after offloading. Second, delayed KV cache writeback hides storage write latency and mitigates storage write amplification. Finally, a memory-efficient attention accelerator sustains high throughput for long sequences within the resource constraints of NSP devices. We implemented and evaluated HILOS on a real system equipped with 16 SmartSSDs. Compared to state-of-the-art offloading-based inference frameworks, HILOS achieves up to 7.86x throughput while reducing energy consumption by up to 85\%. The source code for HILOS is available at https://github.com/hongsunjang/HILOS.
format Preprint
id arxiv_https___arxiv_org_abs_2502_09921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
Jang, Hongsun
Song, Jaeyong
Shin, Changmin
Noh, Si Ung
Jung, Jaewon
Park, Jisung
Lee, Jinho
Hardware Architecture
The computational and memory demands of large language models for generative inference present significant challenges for practical deployment. One promising solution targeting offline inference is offloading-based batched inference, which extends the GPU's memory hierarchy with host memory and storage. However, it often suffers from substantial I/O overhead, primarily due to the large KV cache sizes that scale with batch size and context window length. In this paper, we introduce HILOS, a framework that boosts offline inference throughput using near-storage processing. The core of HILOS is attention near storage, which offloads memory-intensive attention operations to near-storage accelerators, reducing traffic across the system interconnect. Building on attention near storage, HILOS incorporates three additional optimizations. First, cooperative X-cache minimizes KV cache I/O by exploiting available host resources after offloading. Second, delayed KV cache writeback hides storage write latency and mitigates storage write amplification. Finally, a memory-efficient attention accelerator sustains high throughput for long sequences within the resource constraints of NSP devices. We implemented and evaluated HILOS on a real system equipped with 16 SmartSSDs. Compared to state-of-the-art offloading-based inference frameworks, HILOS achieves up to 7.86x throughput while reducing energy consumption by up to 85\%. The source code for HILOS is available at https://github.com/hongsunjang/HILOS.
title A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
topic Hardware Architecture
url https://arxiv.org/abs/2502.09921