Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Yixuan Even, Savani, Yash, Fang, Fei, Kolter, J. Zico
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917426766020608
author Xu, Yixuan Even
Savani, Yash
Fang, Fei
Kolter, J. Zico
author_facet Xu, Yixuan Even
Savani, Yash
Fang, Fei
Kolter, J. Zico
contents Reinforcement learning with verifiable rewards (RLVR) has emerged as the leading approach for enhancing reasoning capabilities in large language models. However, it faces a fundamental compute and memory asymmetry: rollout generation is embarrassingly parallel and memory-light, whereas policy updates are communication-heavy and memory-intensive. To address this, we introduce PODS (Policy Optimization with Down-Sampling), which decouples rollout generation from policy updates by training only on a strategically selected subset of rollouts, maintaining learning quality while dramatically reducing update costs. We propose a principled subset selection criterion, max-variance down-sampling, that maximizes reward diversity, and provide an efficient $O(n\log n)$ implementation. Empirically, Group Relative Policy Optimization (GRPO) with PODS achieves the peak test accuracy of vanilla GRPO at least $\mathbf{1.7\times}$ faster across the different reasoning benchmarks and hardware configurations we tested.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Xu, Yixuan Even
Savani, Yash
Fang, Fei
Kolter, J. Zico
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement learning with verifiable rewards (RLVR) has emerged as the leading approach for enhancing reasoning capabilities in large language models. However, it faces a fundamental compute and memory asymmetry: rollout generation is embarrassingly parallel and memory-light, whereas policy updates are communication-heavy and memory-intensive. To address this, we introduce PODS (Policy Optimization with Down-Sampling), which decouples rollout generation from policy updates by training only on a strategically selected subset of rollouts, maintaining learning quality while dramatically reducing update costs. We propose a principled subset selection criterion, max-variance down-sampling, that maximizes reward diversity, and provide an efficient $O(n\log n)$ implementation. Empirically, Group Relative Policy Optimization (GRPO) with PODS achieves the peak test accuracy of vanilla GRPO at least $\mathbf{1.7\times}$ faster across the different reasoning benchmarks and hardware configurations we tested.
title Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.13818