Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bauer, Justin, Walshe, Thomas, Pham, Derek, Vishwakarma, Harit, Parchami, Armin, Sala, Frederic, Varma, Paroma
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911608348868608
author Bauer, Justin
Walshe, Thomas
Pham, Derek
Vishwakarma, Harit
Parchami, Armin
Sala, Frederic
Varma, Paroma
author_facet Bauer, Justin
Walshe, Thomas
Pham, Derek
Vishwakarma, Harit
Parchami, Armin
Sala, Frederic
Varma, Paroma
contents Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Reinforcement Learning with Verifiable Rewards (RLVR). While previous work has explored the benefits to model reasoning capabilities by scaling both data and compute used for RLVR, these results lack applicability in many real-world settings where annotated data and accessible compute may be scarce. In this work, we present a comprehensive empirical study of open-source Small Language Model (SLM) performance after RLVR in low data regimes. Across three novel datasets covering number counting problems, graph reasoning, and spatial reasoning, we characterize how model performance scales with dataset size, diversity, and complexity. We demonstrate that (1) procedural datasets allow for fine-grained evaluation and training dataset development with controllable properties (size, diversity, and complexity), (2) under RLVR, models trained on lower complexity tasks can generalize to higher complexity tasks, and (3) training on mixed complexity datasets is associated with the greatest benefits in low data regimes, providing up to 5x sample efficiency versus training on easy tasks. These findings inspire future work on the development of data scaling laws for RLVR and the use of procedural data generators to further understand effective data development for efficient LLM fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18381
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
Bauer, Justin
Walshe, Thomas
Pham, Derek
Vishwakarma, Harit
Parchami, Armin
Sala, Frederic
Varma, Paroma
Artificial Intelligence
Machine Learning
Fine-tuning Large Language Models (LLMs) typically relies on large quantities of high-quality annotated data, or questions with well-defined ground truth answers in the case of Reinforcement Learning with Verifiable Rewards (RLVR). While previous work has explored the benefits to model reasoning capabilities by scaling both data and compute used for RLVR, these results lack applicability in many real-world settings where annotated data and accessible compute may be scarce. In this work, we present a comprehensive empirical study of open-source Small Language Model (SLM) performance after RLVR in low data regimes. Across three novel datasets covering number counting problems, graph reasoning, and spatial reasoning, we characterize how model performance scales with dataset size, diversity, and complexity. We demonstrate that (1) procedural datasets allow for fine-grained evaluation and training dataset development with controllable properties (size, diversity, and complexity), (2) under RLVR, models trained on lower complexity tasks can generalize to higher complexity tasks, and (3) training on mixed complexity datasets is associated with the greatest benefits in low data regimes, providing up to 5x sample efficiency versus training on easy tasks. These findings inspire future work on the development of data scaling laws for RLVR and the use of procedural data generators to further understand effective data development for efficient LLM fine-tuning.
title Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.18381