A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agullo, Ferran, Oliveras, Joan, Wang, Chen, Gutierrez-Torre, Alberto, Tardieu, Olivier, Youssef, Alaa, Torres, Jordi, Berral, Josep Ll.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914163883769856
author Agullo, Ferran
Oliveras, Joan
Wang, Chen
Gutierrez-Torre, Alberto
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
author_facet Agullo, Ferran
Oliveras, Joan
Wang, Chen
Gutierrez-Torre, Alberto
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
contents With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds or thousands of these adapters on a single GPU allows request aggregation, increasing throughput, but may also cause request starvation if GPU memory limits are exceeded. To address this issue, this study focuses on determining the joint configuration of concurrent and parallel adapters that maximizes GPU throughput without inducing starvation, given heterogeneous adapter and traffic properties. We propose a data-driven ML approach leveraging interpretable models to tackle this caching problem and introduce the first Digital Twin capable of reproducing an LLM-adapter serving system, enabling efficient training data generation. Experiments with the vLLM framework and LoRA adapters show that the Digital Twin reproduces throughput within 5.1% of real results, while the ML approach predicts optimal numbers of concurrent and parallel adapters with an error of at most 7.2% under heterogeneous, real-world workloads. The code is publicly available at https://github.com/FerranAgulloLopez/GPULLMAdapterOptimization.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
Agullo, Ferran
Oliveras, Joan
Wang, Chen
Gutierrez-Torre, Alberto
Tardieu, Olivier
Youssef, Alaa
Torres, Jordi
Berral, Josep Ll.
Performance
Artificial Intelligence
Computation and Language
With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds or thousands of these adapters on a single GPU allows request aggregation, increasing throughput, but may also cause request starvation if GPU memory limits are exceeded. To address this issue, this study focuses on determining the joint configuration of concurrent and parallel adapters that maximizes GPU throughput without inducing starvation, given heterogeneous adapter and traffic properties. We propose a data-driven ML approach leveraging interpretable models to tackle this caching problem and introduce the first Digital Twin capable of reproducing an LLM-adapter serving system, enabling efficient training data generation. Experiments with the vLLM framework and LoRA adapters show that the Digital Twin reproduces throughput within 5.1% of real results, while the ML approach predicts optimal numbers of concurrent and parallel adapters with an error of at most 7.2% under heterogeneous, real-world workloads. The code is publicly available at https://github.com/FerranAgulloLopez/GPULLMAdapterOptimization.
title A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
topic Performance
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.08343