Saved in:
Bibliographic Details
Main Authors: Mahrooghi, Ilia, Lotfi, Aryo, Abbe, Emmanuel
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.14868
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913100514459648
author Mahrooghi, Ilia
Lotfi, Aryo
Abbe, Emmanuel
author_facet Mahrooghi, Ilia
Lotfi, Aryo
Abbe, Emmanuel
contents Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly sample-inefficient, as models must navigate vast search spaces with minimal feedback. While classic curriculum learning aims to mitigate this by ordering data based on complexity, prior works have primarily targeted small datasets and do not directly transfer to the large-scale settings typical of modern LM training. Furthermore, the right ordering for a specific model is often unclear. To address this, we propose Goldilocks, a novel teacher-driven data sampling strategy that aims to predict each question's difficulty for the student model. The teacher model selects questions of appropriate difficulty for the student model, i.e., questions that are neither too easy nor too hard (Goldilocks principle), while training the student with GRPO. By leveraging the student's performance on seen samples, the teacher continuously adapts to the student's evolving abilities. On the OpenMathReasoning dataset, Goldilocks data sampling improves the performance of models trained with standard GRPO under the same compute budget.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14868
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
Mahrooghi, Ilia
Lotfi, Aryo
Abbe, Emmanuel
Machine Learning
Artificial Intelligence
Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly sample-inefficient, as models must navigate vast search spaces with minimal feedback. While classic curriculum learning aims to mitigate this by ordering data based on complexity, prior works have primarily targeted small datasets and do not directly transfer to the large-scale settings typical of modern LM training. Furthermore, the right ordering for a specific model is often unclear. To address this, we propose Goldilocks, a novel teacher-driven data sampling strategy that aims to predict each question's difficulty for the student model. The teacher model selects questions of appropriate difficulty for the student model, i.e., questions that are neither too easy nor too hard (Goldilocks principle), while training the student with GRPO. By leveraging the student's performance on seen samples, the teacher continuously adapts to the student's evolving abilities. On the OpenMathReasoning dataset, Goldilocks data sampling improves the performance of models trained with standard GRPO under the same compute budget.
title Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.14868