Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qu, Yun, Wang, Qi, Mao, Yixiu, Hu, Vincent Tao, Ommer, Björn, Ji, Xiangyang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918280347779072
author Qu, Yun
Wang, Qi
Mao, Yixiu
Hu, Vincent Tao
Ommer, Björn
Ji, Xiangyang
author_facet Qu, Yun
Wang, Qi
Mao, Yixiu
Hu, Vincent Tao
Ommer, Björn
Ji, Xiangyang
contents Recent advances have witnessed the effectiveness of reinforcement learning (RL) finetuning in enhancing the reasoning capabilities of large language models (LLMs). The optimization process often requires numerous iterations to achieve satisfactory performance, resulting in high computational costs due to the need for frequent prompt evaluations under intensive LLM interactions and repeated policy updates. Appropriate online prompt selection methods reduce iteration steps by prioritizing informative prompts during training, while the pipeline's reliance on exhaustive prompt evaluation and subset selection for optimization still incurs substantial computational overhead due to frequent LLM inference calls. Distinguished from these direct evaluate-then-select schemes, this work investigates iterative approximate evaluation for arbitrary prompts and introduces Model Predictive Prompt Selection (MoPPS), a Bayesian risk-predictive framework that online estimates prompt difficulty without requiring costly LLM interactions. Technically, MoPPS models each prompt's success rate as a latent variable, performs streaming Bayesian inference, and employs posterior sampling in a constructed multi-armed bandit machine, enabling sample efficient and adaptive prompt selection. Extensive experiments across mathematics, planning, and vision-based geometry tasks show that MoPPS reliably predicts prompt difficulty and accelerates training with significantly reduced LLM rollouts. Our code is available at https://github.com/thu-rllab/MoPPS.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
Qu, Yun
Wang, Qi
Mao, Yixiu
Hu, Vincent Tao
Ommer, Björn
Ji, Xiangyang
Artificial Intelligence
Machine Learning
Recent advances have witnessed the effectiveness of reinforcement learning (RL) finetuning in enhancing the reasoning capabilities of large language models (LLMs). The optimization process often requires numerous iterations to achieve satisfactory performance, resulting in high computational costs due to the need for frequent prompt evaluations under intensive LLM interactions and repeated policy updates. Appropriate online prompt selection methods reduce iteration steps by prioritizing informative prompts during training, while the pipeline's reliance on exhaustive prompt evaluation and subset selection for optimization still incurs substantial computational overhead due to frequent LLM inference calls. Distinguished from these direct evaluate-then-select schemes, this work investigates iterative approximate evaluation for arbitrary prompts and introduces Model Predictive Prompt Selection (MoPPS), a Bayesian risk-predictive framework that online estimates prompt difficulty without requiring costly LLM interactions. Technically, MoPPS models each prompt's success rate as a latent variable, performs streaming Bayesian inference, and employs posterior sampling in a constructed multi-armed bandit machine, enabling sample efficient and adaptive prompt selection. Extensive experiments across mathematics, planning, and vision-based geometry tasks show that MoPPS reliably predicts prompt difficulty and accelerates training with significantly reduced LLM rollouts. Our code is available at https://github.com/thu-rllab/MoPPS.
title Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.04632