HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Antian, Zhao, Zhigang, Zhang, Kai, Shi, Xuri, Li, Chuantao, Wang, Chunxiao, He, Zhenying, Jing, Yinan, Wang, X. Sean
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915974674907136
author Liang, Antian
Zhao, Zhigang
Zhang, Kai
Shi, Xuri
Li, Chuantao
Wang, Chunxiao
He, Zhenying
Jing, Yinan
Wang, X. Sean
author_facet Liang, Antian
Zhao, Zhigang
Zhang, Kai
Shi, Xuri
Li, Chuantao
Wang, Chunxiao
He, Zhenying
Jing, Yinan
Wang, X. Sean
contents With the rapid evolution of GPU architectures, the heterogeneity of model training infrastructures is steadily increasing. In such environments, effectively utilizing all available heterogeneous accelerators becomes critical for distributed model training. However, existing frameworks, which are primarily designed for homogeneous clusters, often exhibit significant resource underutilization when deployed on heterogeneous accelerators and networks. In this paper, we present Harp, an automated parallel training framework designed specifically for heterogeneous clusters. Harp introduces a fine-grained planner that efficiently searches a wide space for the inter-operator parallel strategy, enabling Harp to alleviate communication overheads while maintaining balanced loads across heterogeneous accelerators. In addition, Harp implements a heterogeneity-aware 1F1B scheduler that adaptively adjusts the execution timing and ordering of microbatches based on network characteristics, maximizing computation-communication overlap under cross-cluster interconnects while incurring only minimal memory overhead. Our evaluation results show that Harp can deliver 1.3x-1.6x higher performance on heterogeneous clusters than state-of-the-art training frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
Liang, Antian
Zhao, Zhigang
Zhang, Kai
Shi, Xuri
Li, Chuantao
Wang, Chunxiao
He, Zhenying
Jing, Yinan
Wang, X. Sean
Distributed, Parallel, and Cluster Computing
With the rapid evolution of GPU architectures, the heterogeneity of model training infrastructures is steadily increasing. In such environments, effectively utilizing all available heterogeneous accelerators becomes critical for distributed model training. However, existing frameworks, which are primarily designed for homogeneous clusters, often exhibit significant resource underutilization when deployed on heterogeneous accelerators and networks. In this paper, we present Harp, an automated parallel training framework designed specifically for heterogeneous clusters. Harp introduces a fine-grained planner that efficiently searches a wide space for the inter-operator parallel strategy, enabling Harp to alleviate communication overheads while maintaining balanced loads across heterogeneous accelerators. In addition, Harp implements a heterogeneity-aware 1F1B scheduler that adaptively adjusts the execution timing and ordering of microbatches based on network characteristics, maximizing computation-communication overlap under cross-cluster interconnects while incurring only minimal memory overhead. Our evaluation results show that Harp can deliver 1.3x-1.6x higher performance on heterogeneous clusters than state-of-the-art training frameworks.
title HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.24859