Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Yiju, Hu, Tianyi, Sun, Zexu, Lin, Yankai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918454328557568
author Guo, Yiju
Hu, Tianyi
Sun, Zexu
Lin, Yankai
author_facet Guo, Yiju
Hu, Tianyi
Sun, Zexu
Lin, Yankai
contents Reinforcement Learning with Verifiable Rewards (RLVR) has advanced LLM reasoning, but remains constrained by inefficient exploration under limited rollout budgets, leading to low sampling success and unstable training in complex tasks. We find that many exploration failures arise not from problem difficulty, but from a small number of prompt tokens that introduce interference. Building on this insight, we propose the Less Noise Sampling Framework (LENS), which first prompts by identifying and removing interference tokens. then transfers successful rollouts from the purification process to supervise policy optimization on the original noisy prompts, enabling the model to learn to ignore interference in the real-world, noisy prompting settings. Experimental results show that LENS significantly outperforms GRPO, delivering higher performance and faster convergence, with a 3.88% average gain and over 1.6$\times$ speedup on math reasoning, and a 1.83% gain on scientific and general reasoning. Our work highlights the critical role of pruning interference tokens in improving rollout efficiency, offering a new perspective for RLVR research.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21244
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
Guo, Yiju
Hu, Tianyi
Sun, Zexu
Lin, Yankai
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) has advanced LLM reasoning, but remains constrained by inefficient exploration under limited rollout budgets, leading to low sampling success and unstable training in complex tasks. We find that many exploration failures arise not from problem difficulty, but from a small number of prompt tokens that introduce interference. Building on this insight, we propose the Less Noise Sampling Framework (LENS), which first prompts by identifying and removing interference tokens. then transfers successful rollouts from the purification process to supervise policy optimization on the original noisy prompts, enabling the model to learn to ignore interference in the real-world, noisy prompting settings. Experimental results show that LENS significantly outperforms GRPO, delivering higher performance and faster convergence, with a 3.88% average gain and over 1.6$\times$ speedup on math reasoning, and a 1.83% gain on scientific and general reasoning. Our work highlights the critical role of pruning interference tokens in improving rollout efficiency, offering a new perspective for RLVR research.
title Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.21244