Data-Efficient RLVR via Off-Policy Influence Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Erle, Jiang, Dazhi, Wang, Yuan, Li, Xujun, Cheng, Jiale, Gu, Yuxian, Niu, Yilin, Zeng, Aohan, Tang, Jie, Huang, Minlie, Wang, Hongning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915937479819264
author Zhu, Erle
Jiang, Dazhi
Wang, Yuan
Li, Xujun
Cheng, Jiale
Gu, Yuxian
Niu, Yilin
Zeng, Aohan
Tang, Jie
Huang, Minlie
Wang, Hongning
author_facet Zhu, Erle
Jiang, Dazhi
Wang, Yuan
Li, Xujun
Cheng, Jiale
Gu, Yuxian
Niu, Yilin
Zeng, Aohan
Tang, Jie
Huang, Minlie
Wang, Hongning
contents Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large language models (LLMs). Current data selection methods are largely heuristic-based, lacking theoretical guarantees and generalizability. This work proposes a theoretically-grounded approach using influence functions to estimate the contribution of each data point to the learning objective. To overcome the prohibitive computational cost of policy rollouts required for online influence estimation, we introduce an off-policy influence estimation method that efficiently approximates data influence using pre-collected offline trajectories. Furthermore, to manage the high-dimensional gradients of LLMs, we employ sparse random projection to reduce dimensionality and improve storage and computation efficiency. Leveraging these techniques, we develop \textbf{C}urriculum \textbf{R}L with \textbf{O}ff-\textbf{P}olicy \text{I}nfluence guidance (\textbf{CROPI}), a multi-stage RL framework that iteratively selects the most influential data for the current policy. Experiments on models up to 7B parameters demonstrate that CROPI significantly accelerates training. On a 1.5B model, it achieves a 2.66x step-level acceleration while using only 10\% of the data per stage compared to full-dataset training. Our results highlight the substantial potential of influence-based data selection for efficient RLVR.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26491
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data-Efficient RLVR via Off-Policy Influence Guidance
Zhu, Erle
Jiang, Dazhi
Wang, Yuan
Li, Xujun
Cheng, Jiale
Gu, Yuxian
Niu, Yilin
Zeng, Aohan
Tang, Jie
Huang, Minlie
Wang, Hongning
Machine Learning
Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large language models (LLMs). Current data selection methods are largely heuristic-based, lacking theoretical guarantees and generalizability. This work proposes a theoretically-grounded approach using influence functions to estimate the contribution of each data point to the learning objective. To overcome the prohibitive computational cost of policy rollouts required for online influence estimation, we introduce an off-policy influence estimation method that efficiently approximates data influence using pre-collected offline trajectories. Furthermore, to manage the high-dimensional gradients of LLMs, we employ sparse random projection to reduce dimensionality and improve storage and computation efficiency. Leveraging these techniques, we develop \textbf{C}urriculum \textbf{R}L with \textbf{O}ff-\textbf{P}olicy \text{I}nfluence guidance (\textbf{CROPI}), a multi-stage RL framework that iteratively selects the most influential data for the current policy. Experiments on models up to 7B parameters demonstrate that CROPI significantly accelerates training. On a 1.5B model, it achieves a 2.66x step-level acceleration while using only 10\% of the data per stage compared to full-dataset training. Our results highlight the substantial potential of influence-based data selection for efficient RLVR.
title Data-Efficient RLVR via Off-Policy Influence Guidance
topic Machine Learning
url https://arxiv.org/abs/2510.26491