XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
Fuente:
arXiv
Saved in:
| Main Authors: | Bamba, Udbhav, Fang, Minghao, Yu, Yifan, Zheng, Haizhong, Lai, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
by: Yan, Kaizhuo, et al.
Published: (2025)
by: Yan, Kaizhuo, et al.
Published: (2025)
DOT-MoE: Differentiable Optimal Transport for MoEfication
by: Bamba, Udbhav, et al.
Published: (2026)
by: Bamba, Udbhav, et al.
Published: (2026)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
by: Chavan, Arnav, et al.
Published: (2026)
by: Chavan, Arnav, et al.
Published: (2026)
Neural Exploitation and Exploration of Contextual Bandits
by: Ban, Yikun, et al.
Published: (2023)
by: Ban, Yikun, et al.
Published: (2023)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
by: Lai, Yao, et al.
Published: (2026)
by: Lai, Yao, et al.
Published: (2026)
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
by: Yan, Renye, et al.
Published: (2024)
by: Yan, Renye, et al.
Published: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
by: Le, Khiem, et al.
Published: (2026)
by: Le, Khiem, et al.
Published: (2026)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
Compass: SLO-aware Query Planner for Compound AI Serving at Scale
by: Liu, Banruo, et al.
Published: (2025)
by: Liu, Banruo, et al.
Published: (2025)
Exploration-Exploitation Tradeoff in Universal Lossy Compression
by: Weinberger, Nir, et al.
Published: (2025)
by: Weinberger, Nir, et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
by: Li, Zhaochun, et al.
Published: (2026)
by: Li, Zhaochun, et al.
Published: (2026)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Structured Exploration and Exploitation of Label Functions for Automated Data Annotation
by: Lam, Phong, et al.
Published: (2026)
by: Lam, Phong, et al.
Published: (2026)
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
by: Li, Xiaofan, et al.
Published: (2026)
by: Li, Xiaofan, et al.
Published: (2026)
Class-Proportional Coreset Selection for Difficulty-Separable Data
by: Tsai, Elisa, et al.
Published: (2025)
by: Tsai, Elisa, et al.
Published: (2025)
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
by: Zhen, Shuai, et al.
Published: (2026)
by: Zhen, Shuai, et al.
Published: (2026)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Disentangling Exploration of Large Language Models by Optimal Exploitation
by: Grams, Tim, et al.
Published: (2025)
by: Grams, Tim, et al.
Published: (2025)
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
by: Ding, Zheng, et al.
Published: (2025)
by: Ding, Zheng, et al.
Published: (2025)
Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
by: Niu, Zenghao, et al.
Published: (2025)
by: Niu, Zenghao, et al.
Published: (2025)
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
by: Mukherjee, Sumantrak, et al.
Published: (2025)
by: Mukherjee, Sumantrak, et al.
Published: (2025)
Exploitation Over Exploration: Unmasking the Bias in Linear Bandit Recommender Offline Evaluation
by: Pires, Pedro R., et al.
Published: (2025)
by: Pires, Pedro R., et al.
Published: (2025)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
by: Zheng, Haizhong, et al.
Published: (2024)
by: Zheng, Haizhong, et al.
Published: (2024)
Pushing the limits of unconstrained machine-learned interatomic potentials
by: Bigi, Filippo, et al.
Published: (2026)
by: Bigi, Filippo, et al.
Published: (2026)
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
by: Nie, Wenhua, et al.
Published: (2026)
by: Nie, Wenhua, et al.
Published: (2026)
From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
by: Xu, Donglai, et al.
Published: (2025)
by: Xu, Donglai, et al.
Published: (2025)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs
by: Norman, Ben, et al.
Published: (2023)
by: Norman, Ben, et al.
Published: (2023)
Similar Items
-
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
by: Yan, Kaizhuo, et al.
Published: (2025) -
DOT-MoE: Differentiable Optimal Transport for MoEfication
by: Bamba, Udbhav, et al.
Published: (2026) -
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026) -
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
by: Chavan, Arnav, et al.
Published: (2026) -
Neural Exploitation and Exploration of Contextual Bandits
by: Ban, Yikun, et al.
Published: (2023)