JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hongyu, Liu, Weijian, Xu, Hongtao, Wang, Yan, Li, Mingzhen, Jia, Weile, Tan, Guangming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917510918438912
author Wang, Hongyu
Liu, Weijian
Xu, Hongtao
Wang, Yan
Li, Mingzhen
Jia, Weile
Tan, Guangming
author_facet Wang, Hongyu
Liu, Weijian
Xu, Hongtao
Wang, Yan
Li, Mingzhen
Jia, Weile
Tan, Guangming
contents Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scalable and efficient distributed training systems for conservative MLIPs makes them difficult to scale. This is because conservative MLIPs inherently follow a double-backward execution pattern, which involves computing gradients during the forward pass. This pattern creates a mismatch with existing distributed training systems, especially for pipeline parallelism. Therefore, we present JanusPipe, an efficient 3D-parallel (PP/DP/GP) training system tailored for conservative MLIPs. It integrates SymFold to enable memory-efficient pipeline parallelism for conservative MLIPs, and WaveK to reduce pipeline bubbles by balancing the four-phase compute time. Experimental results on 32 GPUs show that JanusPipe improves throughput by $1.51\times$ and $1.45\times$ on average over 1F1B and Hanayo, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18404
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
Wang, Hongyu
Liu, Weijian
Xu, Hongtao
Wang, Yan
Li, Mingzhen
Jia, Weile
Tan, Guangming
Distributed, Parallel, and Cluster Computing
Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scalable and efficient distributed training systems for conservative MLIPs makes them difficult to scale. This is because conservative MLIPs inherently follow a double-backward execution pattern, which involves computing gradients during the forward pass. This pattern creates a mismatch with existing distributed training systems, especially for pipeline parallelism. Therefore, we present JanusPipe, an efficient 3D-parallel (PP/DP/GP) training system tailored for conservative MLIPs. It integrates SymFold to enable memory-efficient pipeline parallelism for conservative MLIPs, and WaveK to reduce pipeline bubbles by balancing the four-phase compute time. Experimental results on 32 GPUs show that JanusPipe improves throughput by $1.51\times$ and $1.45\times$ on average over 1F1B and Hanayo, respectively.
title JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.18404