Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Yun, Wang, Qi, Mao, Yixiu, Zou, Heming, Jiang, Yuhang, Liu, Weijie, Bai, Clive, Yang, Kai, Chen, Yangkun, Yang, Saiyong, Ji, Xiangyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
by: Xu, Xin, et al.
Published: (2026)
by: Xu, Xin, et al.
Published: (2026)
Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
by: Qu, Yun, et al.
Published: (2025)
by: Qu, Yun, et al.
Published: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
Doubly Mild Generalization for Offline Reinforcement Learning
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
by: Tang, Chenming, et al.
Published: (2025)
by: Tang, Chenming, et al.
Published: (2025)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
by: Mao, Yixiu, et al.
Published: (2025)
by: Mao, Yixiu, et al.
Published: (2025)
Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Debiased Model-based Representations for Sample-efficient Continuous Control
by: Lyu, Jiafei, et al.
Published: (2026)
by: Lyu, Jiafei, et al.
Published: (2026)
Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression
by: Mao, Yixiu, et al.
Published: (2024)
by: Mao, Yixiu, et al.
Published: (2024)
Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning
by: Qu, Yun, et al.
Published: (2024)
by: Qu, Yun, et al.
Published: (2024)
Model Predictive Task Sampling for Efficient and Robust Adaptation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments
by: Qu, Yun, et al.
Published: (2025)
by: Qu, Yun, et al.
Published: (2025)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Robust Fast Adaptation from Adversarially Explicit Task Distribution Generation
by: Wang, Cheems, et al.
Published: (2024)
by: Wang, Cheems, et al.
Published: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models
by: Zou, Jade, et al.
Published: (2026)
by: Zou, Jade, et al.
Published: (2026)
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
by: Wu, Jiahao, et al.
Published: (2026)
by: Wu, Jiahao, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
Structural features of the fly olfactory circuit mitigate the stability-plasticity dilemma in continual learning
by: Zou, Heming, et al.
Published: (2025)
by: Zou, Heming, et al.
Published: (2025)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
by: Tang, Chenming, et al.
Published: (2026)
by: Tang, Chenming, et al.
Published: (2026)
Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models
by: Ieong, Lok-Lam, et al.
Published: (2026)
by: Ieong, Lok-Lam, et al.
Published: (2026)
interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification
by: Bhat, Vishak K, et al.
Published: (2026)
by: Bhat, Vishak K, et al.
Published: (2026)
PostCast: Generalizable Postprocessing for Precipitation Nowcasting via Unsupervised Blurriness Modeling
by: Gong, Junchao, et al.
Published: (2024)
by: Gong, Junchao, et al.
Published: (2024)
Learning Dynamics in RL Post-Training for Language Models
by: Tomihari, Akiyoshi
Published: (2026)
by: Tomihari, Akiyoshi
Published: (2026)
Watermarks for Language Models via Probabilistic Automata
by: Wang, Yangkun, et al.
Published: (2025)
by: Wang, Yangkun, et al.
Published: (2025)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training
by: Wang, Chen, et al.
Published: (2026)
by: Wang, Chen, et al.
Published: (2026)
Can Small Language Models Help Large Language Models Reason Better?: LM-Guided Chain-of-Thought
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
by: Sinii, Viacheslav, et al.
Published: (2025)
by: Sinii, Viacheslav, et al.
Published: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
by: Chen, Yangkun, et al.
Published: (2024)
by: Chen, Yangkun, et al.
Published: (2024)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Similar Items
-
Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
by: Mao, Yixiu, et al.
Published: (2026) -
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026) -
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
by: Xu, Xin, et al.
Published: (2026) -
Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
by: Qu, Yun, et al.
Published: (2025) -
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)