Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hang, Shi, Jiuchen, Wang, Yixiao, Chen, Quan, Shan, Yizhou, Guo, Minyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912363220828160
author Zhang, Hang
Shi, Jiuchen
Wang, Yixiao
Chen, Quan
Shan, Yizhou
Guo, Minyi
author_facet Zhang, Hang
Shi, Jiuchen
Wang, Yixiao
Chen, Quan
Shan, Yizhou
Guo, Minyi
contents Multiple Low-Rank Adapters (Multi-LoRAs) are gaining popularity for task-specific Large Language Model (LLM) applications. For multi-LoRA serving, caching hot KV caches and LoRA adapters in high bandwidth memory of accelerations can improve inference performance. However, existing Multi-LoRA inference systems fail to optimize serving performance like Time-To-First-Toke (TTFT), neglecting usage dependencies when caching LoRAs and KVs. We therefore propose FASTLIBRA, a Multi-LoRA caching system to optimize the serving performance. FASTLIBRA comprises a dependency-aware cache manager and a performance-driven cache swapper. The cache manager maintains the usage dependencies between LoRAs and KV caches during the inference with a unified caching pool. The cache swapper determines the swap-in or out of LoRAs and KV caches based on a unified cost model, when the HBM is idle or busy, respectively. Experimental results show that ELORA reduces the TTFT by 63.4% on average, compared to state-of-the-art works.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03756
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Zhang, Hang
Shi, Jiuchen
Wang, Yixiao
Chen, Quan
Shan, Yizhou
Guo, Minyi
Hardware Architecture
Artificial Intelligence
Machine Learning
Performance
Multiple Low-Rank Adapters (Multi-LoRAs) are gaining popularity for task-specific Large Language Model (LLM) applications. For multi-LoRA serving, caching hot KV caches and LoRA adapters in high bandwidth memory of accelerations can improve inference performance. However, existing Multi-LoRA inference systems fail to optimize serving performance like Time-To-First-Toke (TTFT), neglecting usage dependencies when caching LoRAs and KVs. We therefore propose FASTLIBRA, a Multi-LoRA caching system to optimize the serving performance. FASTLIBRA comprises a dependency-aware cache manager and a performance-driven cache swapper. The cache manager maintains the usage dependencies between LoRAs and KV caches during the inference with a unified caching pool. The cache swapper determines the swap-in or out of LoRAs and KV caches based on a unified cost model, when the HBM is idle or busy, respectively. Experimental results show that ELORA reduces the TTFT by 63.4% on average, compared to state-of-the-art works.
title Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
topic Hardware Architecture
Artificial Intelligence
Machine Learning
Performance
url https://arxiv.org/abs/2505.03756