Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Prakhar, Gupta, Vaibhav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026)
by: Salimi, Moein, et al.
Published: (2026)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
by: Ye, Chenlu, et al.
Published: (2025)
by: Ye, Chenlu, et al.
Published: (2025)
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
by: Tang, Pingzhi, et al.
Published: (2026)
by: Tang, Pingzhi, et al.
Published: (2026)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
by: Muslimani, Calarina, et al.
Published: (2025)
by: Muslimani, Calarina, et al.
Published: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
Advancing Safe Mechanical Ventilation Using Offline RL With Hybrid Actions and Clinically Aligned Rewards
by: Yousuf, Muhammad Hamza, et al.
Published: (2025)
by: Yousuf, Muhammad Hamza, et al.
Published: (2025)
A Proposed Paradigm for Imputing Missing Multi-Sensor Data in the Healthcare Domain
by: Gupta, Vaibhav, et al.
Published: (2026)
by: Gupta, Vaibhav, et al.
Published: (2026)
DiRL: An Efficient Post-Training Framework for Diffusion Language Models
by: Zhu, Ying, et al.
Published: (2025)
by: Zhu, Ying, et al.
Published: (2025)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
Policy Learning for Off-Dynamics RL with Deficient Support
by: Van, Linh Le Pham, et al.
Published: (2024)
by: Van, Linh Le Pham, et al.
Published: (2024)
On Designing Effective RL Reward at Training Time for LLM Reasoning
by: Gao, Jiaxuan, et al.
Published: (2024)
by: Gao, Jiaxuan, et al.
Published: (2024)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
by: Han, Zhenyu, et al.
Published: (2025)
by: Han, Zhenyu, et al.
Published: (2025)
Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training
by: Cui, Peng, et al.
Published: (2026)
by: Cui, Peng, et al.
Published: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
by: Chen, Yutong, et al.
Published: (2026)
by: Chen, Yutong, et al.
Published: (2026)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
by: Ikezogwo, Wisdom, et al.
Published: (2026)
by: Ikezogwo, Wisdom, et al.
Published: (2026)
A Training Data Recipe to Accelerate A* Search with Language Models
by: Gupta, Devaansh, et al.
Published: (2024)
by: Gupta, Devaansh, et al.
Published: (2024)
Sampling for Quality: Training-Free Reward-Guided LLM Decoding via Sequential Monte Carlo
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
by: Padula, Alexander G., et al.
Published: (2024)
by: Padula, Alexander G., et al.
Published: (2024)
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
by: Gwak, Minju, et al.
Published: (2026)
by: Gwak, Minju, et al.
Published: (2026)
Self-Improving Tabular Language Models via Iterative Reward-Guided Post-Training
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration
by: Le, Tuan, et al.
Published: (2025)
by: Le, Tuan, et al.
Published: (2025)
Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
by: Nie, Dong
Published: (2026)
by: Nie, Dong
Published: (2026)
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
by: Ren, Tao, et al.
Published: (2025)
by: Ren, Tao, et al.
Published: (2025)
GUI-GENESIS: Automated Synthesis of Efficient Environments with Verifiable Rewards for GUI Agent Post-Training
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Autotelic Reinforcement Learning: Exploring Intrinsic Motivations for Skill Acquisition in Open-Ended Environments
by: Srivastava, Prakhar, et al.
Published: (2025)
by: Srivastava, Prakhar, et al.
Published: (2025)
SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
by: Cuadron, Alejandro, et al.
Published: (2025)
by: Cuadron, Alejandro, et al.
Published: (2025)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
by: Tan, Jinzhou, et al.
Published: (2026)
by: Tan, Jinzhou, et al.
Published: (2026)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
by: Mahrooghi, Ilia, et al.
Published: (2026)
by: Mahrooghi, Ilia, et al.
Published: (2026)
Similar Items
-
Better World Models Can Lead to Better Post-Training Performance
by: Gupta, Prakhar, et al.
Published: (2025) -
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
by: Salimi, Moein, et al.
Published: (2026) -
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
by: Hu, Yuelin, et al.
Published: (2026) -
Bootstrapped Reward Shaping
by: Adamczyk, Jacob, et al.
Published: (2025) -
Self-Mined Hardness for Safety Fine-Tuning
by: Gupta, Prakhar, et al.
Published: (2026)