Saved in:
Bibliographic Details
Main Authors: Brüel-Gabrielsson, Rickard, Zhu, Jiacheng, Bhardwaj, Onkar, Choshen, Leshem, Greenewald, Kristjan, Yurochkin, Mikhail, Solomon, Justin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.00066
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908384503005184
author Brüel-Gabrielsson, Rickard
Zhu, Jiacheng
Bhardwaj, Onkar
Choshen, Leshem
Greenewald, Kristjan
Yurochkin, Mikhail
Solomon, Justin
author_facet Brüel-Gabrielsson, Rickard
Zhu, Jiacheng
Bhardwaj, Onkar
Choshen, Leshem
Greenewald, Kristjan
Yurochkin, Mikhail
Solomon, Justin
contents Fine-tuning large language models (LLMs) with low-rank adaptations (LoRAs) has become common practice, often yielding numerous copies of the same LLM differing only in their LoRA updates. This paradigm presents challenges for systems that serve real-time responses to queries that each involve a different LoRA. Prior works optimize the design of such systems but still require continuous loading and offloading of LoRAs, as it is infeasible to store thousands of LoRAs in GPU memory. To mitigate this issue, we investigate the efficacy of compression when serving LoRAs. We propose a method for the joint compression of LoRAs into a shared basis paired with LoRA-specific scaling matrices. We extend our algorithm to learn clusters of LoRAs that are amenable to joint compression, allowing it to scale gracefully to large LoRA collections. Our experiments with up to 1000 LoRAs demonstrate that compressed LoRAs preserve performance while offering major throughput gains in realistic serving scenarios with over a thousand LoRAs, maintaining 80% of the throughput of serving a single LoRA.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00066
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
Brüel-Gabrielsson, Rickard
Zhu, Jiacheng
Bhardwaj, Onkar
Choshen, Leshem
Greenewald, Kristjan
Yurochkin, Mikhail
Solomon, Justin
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Computation and Language
Machine Learning
Fine-tuning large language models (LLMs) with low-rank adaptations (LoRAs) has become common practice, often yielding numerous copies of the same LLM differing only in their LoRA updates. This paradigm presents challenges for systems that serve real-time responses to queries that each involve a different LoRA. Prior works optimize the design of such systems but still require continuous loading and offloading of LoRAs, as it is infeasible to store thousands of LoRAs in GPU memory. To mitigate this issue, we investigate the efficacy of compression when serving LoRAs. We propose a method for the joint compression of LoRAs into a shared basis paired with LoRA-specific scaling matrices. We extend our algorithm to learn clusters of LoRAs that are amenable to joint compression, allowing it to scale gracefully to large LoRA collections. Our experiments with up to 1000 LoRAs demonstrate that compressed LoRAs preserve performance while offering major throughput gains in realistic serving scenarios with over a thousand LoRAs, maintaining 80% of the throughput of serving a single LoRA.
title Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.00066