DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Xuerui, Guo, Liya, Wang, Yue, Zhu, Yi, Ma, Zhiming, Wang, Zun, Liu, Yuting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
Potential Score Matching: Debiasing Molecular Structure Sampling with Potential Energy Guidance
by: Guo, Liya, et al.
Published: (2025)
by: Guo, Liya, et al.
Published: (2025)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Auto-configuring Exploration-Exploitation Tradeoff in Evolutionary Computation via Deep Reinforcement Learning
by: Ma, Zeyuan, et al.
Published: (2024)
by: Ma, Zeyuan, et al.
Published: (2024)
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning
by: Guo, Weiyang, et al.
Published: (2026)
by: Guo, Weiyang, et al.
Published: (2026)
Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
by: Chen, Lin, et al.
Published: (2026)
by: Chen, Lin, et al.
Published: (2026)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
by: Liu, Jia, et al.
Published: (2025)
by: Liu, Jia, et al.
Published: (2025)
HyPER: Bridging Exploration and Exploitation for Scalable LLM Reasoning with Hypothesis Path Expansion and Reduction
by: Qiu, Shengxuan, et al.
Published: (2026)
by: Qiu, Shengxuan, et al.
Published: (2026)
Balancing Exploration and Exploitation in LLM using Soft RLLF for Enhanced Negation Understanding
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
by: Nguyen, Ha-Thanh, et al.
Published: (2024)
Enhancing LLM Reasoning with Reward-guided Tree Search
by: Jiang, Jinhao, et al.
Published: (2024)
by: Jiang, Jinhao, et al.
Published: (2024)
Plan-MCTS: Plan Exploration for Action Exploitation in Web Navigation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
by: Chen, Zhipeng, et al.
Published: (2025)
by: Chen, Zhipeng, et al.
Published: (2025)
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
by: Bi, Baolong, et al.
Published: (2025)
by: Bi, Baolong, et al.
Published: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
by: Zhang, Qifan, et al.
Published: (2026)
by: Zhang, Qifan, et al.
Published: (2026)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
by: Chen, Sirui, et al.
Published: (2026)
by: Chen, Sirui, et al.
Published: (2026)
Requesting Expert Reasoning: Augmenting LLM Agents with Learned Collaborative Intervention
by: Wang, Zhiming, et al.
Published: (2026)
by: Wang, Zhiming, et al.
Published: (2026)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation
by: Yang, Lu, et al.
Published: (2026)
by: Yang, Lu, et al.
Published: (2026)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
by: Lin, Wenze, et al.
Published: (2026)
by: Lin, Wenze, et al.
Published: (2026)
Efficient Reasoning via Reward Model
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
by: Alahmadi, Muhammad J., et al.
Published: (2026)
by: Alahmadi, Muhammad J., et al.
Published: (2026)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
by: Niu, Zenghao, et al.
Published: (2025)
by: Niu, Zenghao, et al.
Published: (2025)
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
by: Feng, Zhaopeng, et al.
Published: (2025)
by: Feng, Zhaopeng, et al.
Published: (2025)
Enhancing Table Reasoning with Deterministic Table-State Rewards
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
On Designing Effective RL Reward at Training Time for LLM Reasoning
by: Gao, Jiaxuan, et al.
Published: (2024)
by: Gao, Jiaxuan, et al.
Published: (2024)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
Query Decomposition for RAG: Balancing Exploration-Exploitation
by: Petcu, Roxana, et al.
Published: (2025)
by: Petcu, Roxana, et al.
Published: (2025)
VASR: Variance-Aware Systematic Resampling for Reward-Guided Diffusion
by: Shekhar, Shivanshu, et al.
Published: (2026)
by: Shekhar, Shivanshu, et al.
Published: (2026)
Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning
by: Gao, Junqi, et al.
Published: (2025)
by: Gao, Junqi, et al.
Published: (2025)
Similar Items
-
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
by: Su, Xuerui, et al.
Published: (2025) -
Potential Score Matching: Debiasing Molecular Structure Sampling with Potential Energy Guidance
by: Guo, Liya, et al.
Published: (2025) -
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026) -
WESE: Weak Exploration to Strong Exploitation for LLM Agents
by: Huang, Xu, et al.
Published: (2024) -
Auto-configuring Exploration-Exploitation Tradeoff in Evolutionary Computation via Deep Reinforcement Learning
by: Ma, Zeyuan, et al.
Published: (2024)