Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Runsheng Benson, Anand, Utkarsh, Daudjee, Khuzaima, Sen, Rathijit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916843117084672
author Guo, Runsheng Benson
Anand, Utkarsh
Daudjee, Khuzaima
Sen, Rathijit
author_facet Guo, Runsheng Benson
Anand, Utkarsh
Daudjee, Khuzaima
Sen, Rathijit
contents Large language models (LLMs) require vast amounts of GPU compute to train, but limited availability and high costs of GPUs make homogeneous clusters impractical for many organizations. Instead, assembling heterogeneous clusters by pooling together GPUs of different generations allows them to achieve higher aggregate compute and make use of all available GPUs. However, training on heterogeneous clusters presents several challenges, including load balancing across GPUs, optimizing memory usage to accommodate varying memory capacities, and ensuring communication-efficient training over diverse network interconnects potentially spanning multiple datacenters. In this paper, we make the case that efficient training on heterogeneous clusters requires (1) the integration of pipeline parallelism and data parallelism in a manner that is both communication- and memory-efficient, and (2) a more adaptable configuration of pipeline and data parallelism, which includes the capability to flexibly partition GPUs into asymmetric pipeline parallel stages and to incorporate heterogeneous GPUs within the same data parallelism group. We propose Zorse, the first system to unify all these capabilities while incorporating a planner that automatically configures training strategies for a given workload. Our evaluation shows that Zorse significantly outperforms state-of-the-art systems in heterogeneous training scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10392
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
Guo, Runsheng Benson
Anand, Utkarsh
Daudjee, Khuzaima
Sen, Rathijit
Distributed, Parallel, and Cluster Computing
Large language models (LLMs) require vast amounts of GPU compute to train, but limited availability and high costs of GPUs make homogeneous clusters impractical for many organizations. Instead, assembling heterogeneous clusters by pooling together GPUs of different generations allows them to achieve higher aggregate compute and make use of all available GPUs. However, training on heterogeneous clusters presents several challenges, including load balancing across GPUs, optimizing memory usage to accommodate varying memory capacities, and ensuring communication-efficient training over diverse network interconnects potentially spanning multiple datacenters. In this paper, we make the case that efficient training on heterogeneous clusters requires (1) the integration of pipeline parallelism and data parallelism in a manner that is both communication- and memory-efficient, and (2) a more adaptable configuration of pipeline and data parallelism, which includes the capability to flexibly partition GPUs into asymmetric pipeline parallel stages and to incorporate heterogeneous GPUs within the same data parallelism group. We propose Zorse, the first system to unify all these capabilities while incorporating a planner that automatically configures training strategies for a given workload. Our evaluation shows that Zorse significantly outperforms state-of-the-art systems in heterogeneous training scenarios.
title Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2507.10392