DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yao, Xiaozhe, Hu, Qinghao, Klimovic, Ana
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915212550995968
author Yao, Xiaozhe
Hu, Qinghao
Klimovic, Ana
author_facet Yao, Xiaozhe
Hu, Qinghao
Klimovic, Ana
contents Fine-tuning large language models (LLMs) greatly improves model quality for downstream tasks. However, serving many fine-tuned LLMs concurrently is challenging due to the sporadic, bursty, and varying request patterns of different LLMs. To bridge this gap, we present DeltaZip, an LLM serving system that efficiently serves multiple full-parameter fine-tuned models concurrently by aggressively compressing model deltas by up to 10x while maintaining high model quality. The key insight behind this design is that fine-tuning results in small-magnitude changes to the pre-trained model. By co-designing the serving system with the compression algorithm, DeltaZip achieves 2x to 12x improvement in throughput compared to the state-of-the-art systems.
format Preprint
id arxiv_https___arxiv_org_abs_2312_05215
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
Yao, Xiaozhe
Hu, Qinghao
Klimovic, Ana
Distributed, Parallel, and Cluster Computing
Machine Learning
Fine-tuning large language models (LLMs) greatly improves model quality for downstream tasks. However, serving many fine-tuned LLMs concurrently is challenging due to the sporadic, bursty, and varying request patterns of different LLMs. To bridge this gap, we present DeltaZip, an LLM serving system that efficiently serves multiple full-parameter fine-tuned models concurrently by aggressively compressing model deltas by up to 10x while maintaining high model quality. The key insight behind this design is that fine-tuning results in small-magnitude changes to the pre-trained model. By co-designing the serving system with the compression algorithm, DeltaZip achieves 2x to 12x improvement in throughput compared to the state-of-the-art systems.
title DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2312.05215