JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zhengding, Ouyang, Hehua, Chen, Chang, Pan, Zaifeng, Guan, Yue, Yu, Zhongkai, Wang, Zhen, Swanson, Steven, Ding, Yufei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917437692182528
author Hu, Zhengding
Ouyang, Hehua
Chen, Chang
Pan, Zaifeng
Guan, Yue
Yu, Zhongkai
Wang, Zhen
Swanson, Steven
Ding, Yufei
author_facet Hu, Zhengding
Ouyang, Hehua
Chen, Chang
Pan, Zaifeng
Guan, Yue
Yu, Zhongkai
Wang, Zhen
Swanson, Steven
Ding, Yufei
contents We present JigsawRL, a cost-efficient framework that explores Pipeline Multiplexing as a new dimension of RL parallelism. JigsawRL decomposes each pipeline into a Sub-Stage Graph that exposes the intra-stage and inter-worker imbalance hidden by stage-level systems. On this abstraction, JigsawRL resolves multiplexing interference through dynamic resource allocation, eliminates fragmented utilization by migrating long-tail rollouts across workers, and formulates their coordination as a graph scheduling problem solved with a look-ahead heuristic. On 4-64 H100/A100 GPUs across different agentic RL pipelines and models, JigsawRL achieves up to 1.85x throughput over Verl on synchronous RL, 1.54x over StreamRL and AReaL on asynchronous RL, and supports heterogeneous pipelines with moderate latency trade-off.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23838
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
Hu, Zhengding
Ouyang, Hehua
Chen, Chang
Pan, Zaifeng
Guan, Yue
Yu, Zhongkai
Wang, Zhen
Swanson, Steven
Ding, Yufei
Machine Learning
We present JigsawRL, a cost-efficient framework that explores Pipeline Multiplexing as a new dimension of RL parallelism. JigsawRL decomposes each pipeline into a Sub-Stage Graph that exposes the intra-stage and inter-worker imbalance hidden by stage-level systems. On this abstraction, JigsawRL resolves multiplexing interference through dynamic resource allocation, eliminates fragmented utilization by migrating long-tail rollouts across workers, and formulates their coordination as a graph scheduling problem solved with a look-ahead heuristic. On 4-64 H100/A100 GPUs across different agentic RL pipelines and models, JigsawRL achieves up to 1.85x throughput over Verl on synchronous RL, 1.54x over StreamRL and AReaL on asynchronous RL, and supports heterogeneous pipelines with moderate latency trade-off.
title JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
topic Machine Learning
url https://arxiv.org/abs/2604.23838