CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zongkai, Meng, Fanqing, Du, Lingxiao, Zhou, Zhixiang, Yu, Chao, Shao, Wenqi, Zhang, Qiaosheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
by: Du, Lingxiao, et al.
Published: (2025)
by: Du, Lingxiao, et al.
Published: (2025)
Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
by: Zhou, Zhijian, et al.
Published: (2025)
by: Zhou, Zhijian, et al.
Published: (2025)
Policy-regularized Offline Multi-objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024)
by: Lin, Qian, et al.
Published: (2024)
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024)
by: Lin, Qian, et al.
Published: (2024)
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
by: Meng, Fanqing, et al.
Published: (2025)
by: Meng, Fanqing, et al.
Published: (2025)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
by: Liu, Zongkai, et al.
Published: (2024)
by: Liu, Zongkai, et al.
Published: (2024)
Dynamic Adversarial Reinforcement Learning for Robust Multimodal Large Language Models
by: Bao, Yicheng, et al.
Published: (2026)
by: Bao, Yicheng, et al.
Published: (2026)
Ensemble Successor Representations for Task Generalization in Offline-to-Online Reinforcement Learning
by: Wang, Changhong, et al.
Published: (2024)
by: Wang, Changhong, et al.
Published: (2024)
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
by: Shao, Wenqi, et al.
Published: (2023)
by: Shao, Wenqi, et al.
Published: (2023)
Reinforcement Learning-based Feature Generation Algorithm for Scientific Data
by: Xiao, Meng, et al.
Published: (2025)
by: Xiao, Meng, et al.
Published: (2025)
Continual Reinforcement Learning by Planning with Online World Models
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
Enhancing Molecular Property Predictions by Learning from Bond Modelling and Interactions
by: Liu, Yunqing, et al.
Published: (2026)
by: Liu, Yunqing, et al.
Published: (2026)
Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
by: Mathew, Christo, et al.
Published: (2025)
by: Mathew, Christo, et al.
Published: (2025)
Multi-Objective Reinforcement Learning for Large Language Model Optimization: Visionary Perspective
by: Kong, Lingxiao, et al.
Published: (2025)
by: Kong, Lingxiao, et al.
Published: (2025)
$\mathrm{E^{2}CFD}$: Towards Effective and Efficient Cost Function Design for Safe Reinforcement Learning via Large Language Model
by: Wang, Zepeng, et al.
Published: (2024)
by: Wang, Zepeng, et al.
Published: (2024)
The Role of Diversity in In-Context Learning for Large Language Models
by: Xiao, Wenyang, et al.
Published: (2025)
by: Xiao, Wenyang, et al.
Published: (2025)
Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized Worlds
by: Wang, Fan, et al.
Published: (2025)
by: Wang, Fan, et al.
Published: (2025)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Stable Reinforcement Learning for Efficient Reasoning
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Rule-Bottleneck Reinforcement Learning: Joint Explanation and Decision Optimization for Resource Allocation with Language Agents
by: Tec, Mauricio, et al.
Published: (2025)
by: Tec, Mauricio, et al.
Published: (2025)
Towards General Continuous Memory for Vision-Language Models
by: Wu, Wenyi, et al.
Published: (2025)
by: Wu, Wenyi, et al.
Published: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026)
by: Yu, Yongcan, et al.
Published: (2026)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
by: Li, Zhaoxin, et al.
Published: (2025)
by: Li, Zhaoxin, et al.
Published: (2025)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
by: Liu, Max, et al.
Published: (2024)
by: Liu, Max, et al.
Published: (2024)
Continuous Autoregressive Language Models
by: Shao, Chenze, et al.
Published: (2025)
by: Shao, Chenze, et al.
Published: (2025)
Interactive Symbolic Regression through Offline Reinforcement Learning: A Co-Design Framework
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
by: Xie, Sean, et al.
Published: (2022)
by: Xie, Sean, et al.
Published: (2022)
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
by: Xu, Zelai, et al.
Published: (2023)
by: Xu, Zelai, et al.
Published: (2023)
VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Learning Mathematical Rules with Large Language Models
by: Gorceix, Antoine, et al.
Published: (2024)
by: Gorceix, Antoine, et al.
Published: (2024)
Automated discovery of symbolic laws governing skill acquisition from naturally occurring data
by: Liu, Sannyuya, et al.
Published: (2024)
by: Liu, Sannyuya, et al.
Published: (2024)
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
by: Zhang, Weiliang, et al.
Published: (2025)
by: Zhang, Weiliang, et al.
Published: (2025)
Asymptotically Stable Quaternion-valued Hopfield-structured Neural Network with Periodic Projection-based Supervised Learning Rules
by: Wang, Tianwei, et al.
Published: (2025)
by: Wang, Tianwei, et al.
Published: (2025)
The Role of Deductive and Inductive Reasoning in Large Language Models
by: Cai, Chengkun, et al.
Published: (2024)
by: Cai, Chengkun, et al.
Published: (2024)
Towards Differentially Private Reinforcement Learning with General Function Approximation
by: He, Yi, et al.
Published: (2026)
by: He, Yi, et al.
Published: (2026)
Behavior Injection: Preparing Language Models for Reinforcement Learning
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
by: Liu, Ziru, et al.
Published: (2025)
by: Liu, Ziru, et al.
Published: (2025)
Similar Items
-
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
by: Du, Lingxiao, et al.
Published: (2025) -
Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
by: Zhou, Zhijian, et al.
Published: (2025) -
Policy-regularized Offline Multi-objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024) -
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024) -
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
by: Meng, Fanqing, et al.
Published: (2025)