ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Jingwei, Feng, Xinze, Liu, Zien, Wang, Kaijian, Ye, Fanjiang, Cao, Ye, Wang, Zhuang, Wang, Yuke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914462685986816
author Zuo, Jingwei
Feng, Xinze
Liu, Zien
Wang, Kaijian
Ye, Fanjiang
Cao, Ye
Wang, Zhuang
Wang, Yuke
author_facet Zuo, Jingwei
Feng, Xinze
Liu, Zien
Wang, Kaijian
Ye, Fanjiang
Cao, Ye
Wang, Zhuang
Wang, Yuke
contents Low-Rank Adaptation (LoRA) is now the dominant method for parameter-efficient fine-tuning of large language models, but achieving a high-quality adapter often requires systematic hyperparameter tuning because LoRA performance is highly sensitive to configuration choices. In practice, this leads to many concurrent LoRA jobs, often spanning heterogeneous tasks in multi-tenant environments. Existing systems largely handle these jobs independently, which both wastes computation on weak candidates and leaves GPUs underutilized. We present ALTO (Adaptive LoRA Tuning and Orchestration), a co-designed training system that accelerates LoRA hyperparameter tuning while enabling efficient cluster sharing across heterogeneous tasks. The central insight behind ALTO is that when multiple tuning jobs run concurrently over a shared frozen backbone, they expose optimization opportunities that single-job designs cannot exploit. Building on this, ALTO monitors loss trajectories to terminate unpromising configurations early, uses fused grouped GEMM together with a new rank-local adapter parallelism to co-locate surviving adapters and reclaim freed GPU capacity, and combines intra-task and inter-task scheduling to improve multi-task placement by leveraging the predictable duration of LoRA jobs. Extensive evaluation shows that ALTO achieves up to $13.8\times$ speedup over state-of-the-art without sacrificing adapter quality.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05426
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
Zuo, Jingwei
Feng, Xinze
Liu, Zien
Wang, Kaijian
Ye, Fanjiang
Cao, Ye
Wang, Zhuang
Wang, Yuke
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Low-Rank Adaptation (LoRA) is now the dominant method for parameter-efficient fine-tuning of large language models, but achieving a high-quality adapter often requires systematic hyperparameter tuning because LoRA performance is highly sensitive to configuration choices. In practice, this leads to many concurrent LoRA jobs, often spanning heterogeneous tasks in multi-tenant environments. Existing systems largely handle these jobs independently, which both wastes computation on weak candidates and leaves GPUs underutilized. We present ALTO (Adaptive LoRA Tuning and Orchestration), a co-designed training system that accelerates LoRA hyperparameter tuning while enabling efficient cluster sharing across heterogeneous tasks. The central insight behind ALTO is that when multiple tuning jobs run concurrently over a shared frozen backbone, they expose optimization opportunities that single-job designs cannot exploit. Building on this, ALTO monitors loss trajectories to terminate unpromising configurations early, uses fused grouped GEMM together with a new rank-local adapter parallelism to co-locate surviving adapters and reclaim freed GPU capacity, and combines intra-task and inter-task scheduling to improve multi-task placement by leveraging the predictable duration of LoRA jobs. Extensive evaluation shows that ALTO achieves up to $13.8\times$ speedup over state-of-the-art without sacrificing adapter quality.
title ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2604.05426