Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Kai, Schwing, Alexander G., Wang, Yu-Xiong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914567598112768
author Yan, Kai
Schwing, Alexander G.
Wang, Yu-Xiong
author_facet Yan, Kai
Schwing, Alexander G.
Wang, Yu-Xiong
contents Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts for many tasks such as math and coding. Nevertheless, RLVR struggles with sample efficiency on difficult problems where correct rollouts are hard to generate. Prior works propose to address this issue via demonstration-guided RLVR, i.e., to conduct Supervised FineTuning (SFT) when RL fails; however, SFT often requires a lot of data, which can be expensive to acquire. In this paper, we propose FEST, a FEw-ShoT demonstration-guided RLVR algorithm. It attains compelling results with only 128 demonstrations randomly selected from an SFT dataset. We find that three components are vital for the success: supervised signal, on-policy signal, and decaying weights on the few-shot SFT dataset to prevent overfitting from multiple-epoch training. On several benchmarks, FEST outperforms baselines with magnitudes less SFT data, even matching their performance with full dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15012
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
Yan, Kai
Schwing, Alexander G.
Wang, Yu-Xiong
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought rollouts for many tasks such as math and coding. Nevertheless, RLVR struggles with sample efficiency on difficult problems where correct rollouts are hard to generate. Prior works propose to address this issue via demonstration-guided RLVR, i.e., to conduct Supervised FineTuning (SFT) when RL fails; however, SFT often requires a lot of data, which can be expensive to acquire. In this paper, we propose FEST, a FEw-ShoT demonstration-guided RLVR algorithm. It attains compelling results with only 128 demonstrations randomly selected from an SFT dataset. We find that three components are vital for the success: supervised signal, on-policy signal, and decaying weights on the few-shot SFT dataset to prevent overfitting from multiple-epoch training. On several benchmarks, FEST outperforms baselines with magnitudes less SFT data, even matching their performance with full dataset.
title Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.15012