Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Yiming, Xu, Yiran, Lin, Zicheng, Shi, Chufan, Chen, Yukang, Wang, Dingdong, Wu, Tianhe, Wang, Junjie, Yang, Yujiu, Qiao, Yu, Chu, Ruihang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913172160512000
author Ren, Yiming
Xu, Yiran
Lin, Zicheng
Shi, Chufan
Chen, Yukang
Wang, Dingdong
Wu, Tianhe
Wang, Junjie
Yang, Yujiu
Qiao, Yu
Chu, Ruihang
author_facet Ren, Yiming
Xu, Yiran
Lin, Zicheng
Shi, Chufan
Chen, Yukang
Wang, Dingdong
Wu, Tianhe
Wang, Junjie
Yang, Yujiu
Qiao, Yu
Chu, Ruihang
contents We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and lead to incoherent trajectories. We uncover that smaller models within the same model family inherently exhibit higher policy-level diversity, indicated by their superior pass@k relative to larger counterparts as sample counts increase. Unlike token-level noise, this diversity is temporally correlated, preserves logical consistency, and provides structured exploration signals for gradient estimation. We thus propose S2L-PO (Small-to-Large Policy Optimization), a framework that leverages fixed small models as natural explorers to train larger models. To balance exploration and exploitation, we design a progressive annealing strategy that transitions from offline small-model rollouts to the large learner's own sampling. This shift elegantly avoids mid-training performance drops caused by the small model's capacity limits, achieving faster convergence and unlocking a higher performance ceiling. S2L-PO improves accuracy on diverse mathematical reasoning benchmarks (e.g., +8.8% on AIME 24 using a 1.7B explorer to guide the 8B model) while reducing rollout compute.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Ren, Yiming
Xu, Yiran
Lin, Zicheng
Shi, Chufan
Chen, Yukang
Wang, Dingdong
Wu, Tianhe
Wang, Junjie
Yang, Yujiu
Qiao, Yu
Chu, Ruihang
Machine Learning
Artificial Intelligence
We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollouts, prevailing strategies primarily increase diversity by injecting more token-level randomness, which may introduce step-wise noise and lead to incoherent trajectories. We uncover that smaller models within the same model family inherently exhibit higher policy-level diversity, indicated by their superior pass@k relative to larger counterparts as sample counts increase. Unlike token-level noise, this diversity is temporally correlated, preserves logical consistency, and provides structured exploration signals for gradient estimation. We thus propose S2L-PO (Small-to-Large Policy Optimization), a framework that leverages fixed small models as natural explorers to train larger models. To balance exploration and exploitation, we design a progressive annealing strategy that transitions from offline small-model rollouts to the large learner's own sampling. This shift elegantly avoids mid-training performance drops caused by the small model's capacity limits, achieving faster convergence and unlocking a higher performance ceiling. S2L-PO improves accuracy on diverse mathematical reasoning benchmarks (e.g., +8.8% on AIME 24 using a 1.7B explorer to guide the 8B model) while reducing rollout compute.
title Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.30789