Selective Off-Policy Reference Tuning with Plan Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Duc Anh, Nguyen, Tien-Phat, Nguyen, Thien Huu, Van, Linh Ngo, Le, Trung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914562464284672
author Le, Duc Anh
Nguyen, Tien-Phat
Nguyen, Thien Huu
Van, Linh Ngo
Le, Trung
author_facet Le, Duc Anh
Nguyen, Tien-Phat
Nguyen, Thien Huu
Van, Linh Ngo
Le, Trung
contents Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_11505
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Selective Off-Policy Reference Tuning with Plan Guidance
Le, Duc Anh
Nguyen, Tien-Phat
Nguyen, Thien Huu
Van, Linh Ngo
Le, Trung
Artificial Intelligence
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models.
title Selective Off-Policy Reference Tuning with Plan Guidance
topic Artificial Intelligence
url https://arxiv.org/abs/2605.11505