A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914163883769856 |
|---|---|
| author | Agullo, Ferran Oliveras, Joan Wang, Chen Gutierrez-Torre, Alberto Tardieu, Olivier Youssef, Alaa Torres, Jordi Berral, Josep Ll. |
| author_facet | Agullo, Ferran Oliveras, Joan Wang, Chen Gutierrez-Torre, Alberto Tardieu, Olivier Youssef, Alaa Torres, Jordi Berral, Josep Ll. |
| contents | With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds or thousands of these adapters on a single GPU allows request aggregation, increasing throughput, but may also cause request starvation if GPU memory limits are exceeded. To address this issue, this study focuses on determining the joint configuration of concurrent and parallel adapters that maximizes GPU throughput without inducing starvation, given heterogeneous adapter and traffic properties. We propose a data-driven ML approach leveraging interpretable models to tackle this caching problem and introduce the first Digital Twin capable of reproducing an LLM-adapter serving system, enabling efficient training data generation. Experiments with the vLLM framework and LoRA adapters show that the Digital Twin reproduces throughput within 5.1% of real results, while the ML approach predicts optimal numbers of concurrent and parallel adapters with an error of at most 7.2% under heterogeneous, real-world workloads. The code is publicly available at https://github.com/FerranAgulloLopez/GPULLMAdapterOptimization. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_08343 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving Agullo, Ferran Oliveras, Joan Wang, Chen Gutierrez-Torre, Alberto Tardieu, Olivier Youssef, Alaa Torres, Jordi Berral, Josep Ll. Performance Artificial Intelligence Computation and Language With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds or thousands of these adapters on a single GPU allows request aggregation, increasing throughput, but may also cause request starvation if GPU memory limits are exceeded. To address this issue, this study focuses on determining the joint configuration of concurrent and parallel adapters that maximizes GPU throughput without inducing starvation, given heterogeneous adapter and traffic properties. We propose a data-driven ML approach leveraging interpretable models to tackle this caching problem and introduce the first Digital Twin capable of reproducing an LLM-adapter serving system, enabling efficient training data generation. Experiments with the vLLM framework and LoRA adapters show that the Digital Twin reproduces throughput within 5.1% of real results, while the ML approach predicts optimal numbers of concurrent and parallel adapters with an error of at most 7.2% under heterogeneous, real-world workloads. The code is publicly available at https://github.com/FerranAgulloLopez/GPULLMAdapterOptimization. |
| title | A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving |
| topic | Performance Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2508.08343 |