ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910256209068032 |
|---|---|
| author | Song, Jingwei Chen, Meng Xiao, Jie Ren, Qingnan Huang, Jiaqi Deng, Yangshen Tong, Chris Chen, Wanyi Wang, Suli Chen, Zhisheng Bi, Ziqian Lu, Shuo Duan, Yiqun Wang, Xu Yu, Rymon Ai, Lynn Yang, Eric Shi, Tianyu |
| author_facet | Song, Jingwei Chen, Meng Xiao, Jie Ren, Qingnan Huang, Jiaqi Deng, Yangshen Tong, Chris Chen, Wanyi Wang, Suli Chen, Zhisheng Bi, Ziqian Lu, Shuo Duan, Yiqun Wang, Xu Yu, Rymon Ai, Lynn Yang, Eric Shi, Tianyu |
| contents | Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and centralized learning. Distributing rollout execution offers opportunities to leverage more cost-efficient inference resources, but introduces challenges in wide-area coordination and policy dissemination. We present ECHO-2, a distributed RL framework for post-training with remote inference workers and non-negligible dissemination latency. ECHO-2 combines centralized learning with distributed rollouts and treats bounded policy staleness as a user-controlled parameter, enabling rollout generation, dissemination, and training to overlap. We introduce an overlap-based capacity model that relates training time, dissemination latency, and rollout throughput, yielding a practical provisioning rule for sustaining learner utilization. To mitigate dissemination bottlenecks and lower cost, ECHO-2 employs peer-assisted pipelined broadcast and cost-aware activation of heterogeneous workers. Experiments on GRPO post-training of LLMs ranging from 4B to 32B parameters under real wide-area bandwidth regimes show that ECHO-2 significantly improves cost efficiency while preserving RL reward comparable to strong baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_02192 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning Song, Jingwei Chen, Meng Xiao, Jie Ren, Qingnan Huang, Jiaqi Deng, Yangshen Tong, Chris Chen, Wanyi Wang, Suli Chen, Zhisheng Bi, Ziqian Lu, Shuo Duan, Yiqun Wang, Xu Yu, Rymon Ai, Lynn Yang, Eric Shi, Tianyu Machine Learning Distributed, Parallel, and Cluster Computing Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and centralized learning. Distributing rollout execution offers opportunities to leverage more cost-efficient inference resources, but introduces challenges in wide-area coordination and policy dissemination. We present ECHO-2, a distributed RL framework for post-training with remote inference workers and non-negligible dissemination latency. ECHO-2 combines centralized learning with distributed rollouts and treats bounded policy staleness as a user-controlled parameter, enabling rollout generation, dissemination, and training to overlap. We introduce an overlap-based capacity model that relates training time, dissemination latency, and rollout throughput, yielding a practical provisioning rule for sustaining learner utilization. To mitigate dissemination bottlenecks and lower cost, ECHO-2 employs peer-assisted pipelined broadcast and cost-aware activation of heterogeneous workers. Experiments on GRPO post-training of LLMs ranging from 4B to 32B parameters under real wide-area bandwidth regimes show that ECHO-2 significantly improves cost efficiency while preserving RL reward comparable to strong baselines. |
| title | ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2602.02192 |