Selective Off-Policy Reference Tuning with Plan Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914562464284672 |
|---|---|
| author | Le, Duc Anh Nguyen, Tien-Phat Nguyen, Thien Huu Van, Linh Ngo Le, Trung |
| author_facet | Le, Duc Anh Nguyen, Tien-Phat Nguyen, Thien Huu Van, Linh Ngo Le, Trung |
| contents | Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11505 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Selective Off-Policy Reference Tuning with Plan Guidance Le, Duc Anh Nguyen, Tien-Phat Nguyen, Thien Huu Van, Linh Ngo Le, Trung Artificial Intelligence Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones and eight reasoning benchmarks, SORT improves over GRPO and guidance baselines, with largest gains on weaker models. |
| title | Selective Off-Policy Reference Tuning with Plan Guidance |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2605.11505 |