Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Prakhar, Gupta, Vaibhav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910190084816896
author Gupta, Prakhar
Gupta, Vaibhav
author_facet Gupta, Prakhar
Gupta, Vaibhav
contents Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We ask whether a scalar hint toward a canonical solver ordering, used only during RL post-training, improves performance even when fine-tuned on randomized solution sequences. On Zebra puzzles, we fine-tune a Transformer on randomized solution orders, then post-train it with Group Relative Policy Optimization (GRPO) using two rewards: a sparse task reward that is 1 only when the puzzle is fully solved, and an ordering reward that increases when the model's emission order aligns with the canonical solver order. To compare signals cleanly, we combine them via fixed mixtures and use a simple bootstrapped scaling to equalize component magnitudes at initialization. Mixed rewards generally outperform task-only optimization, suggesting that coarse ordering signals can steer RL post-training toward canonical trajectories without modifying supervised data or architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
Gupta, Prakhar
Gupta, Vaibhav
Machine Learning
Artificial Intelligence
Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We ask whether a scalar hint toward a canonical solver ordering, used only during RL post-training, improves performance even when fine-tuned on randomized solution sequences. On Zebra puzzles, we fine-tune a Transformer on randomized solution orders, then post-train it with Group Relative Policy Optimization (GRPO) using two rewards: a sparse task reward that is 1 only when the puzzle is fully solved, and an ordering reward that increases when the model's emission order aligns with the canonical solver order. To compare signals cleanly, we combine them via fixed mixtures and use a simple bootstrapped scaling to equalize component magnitudes at initialization. Mixed rewards generally outperform task-only optimization, suggesting that coarse ordering signals can steer RL post-training toward canonical trajectories without modifying supervised data or architecture.
title Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.04277