Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Wenlong, Ren, Yi, Li, Yushu, Gong, Boying, Sutherland, Danica J., Li, Xiaoxiao, Thrampoulidis, Christos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)
by: Deng, Wenlong, et al.
Published: (2023)
Implicit Optimization Bias of Next-Token Prediction in Linear Models
by: Thrampoulidis, Christos
Published: (2024)
by: Thrampoulidis, Christos
Published: (2024)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
by: Deng, Wenlong, et al.
Published: (2024)
by: Deng, Wenlong, et al.
Published: (2024)
Geometry of Semantics in Next-Token Prediction: How Optimization Implicitly Organizes Linguistic Representations
by: Zhao, Yize, et al.
Published: (2025)
by: Zhao, Yize, et al.
Published: (2025)
Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
by: Huang, Guanhua, et al.
Published: (2025)
by: Huang, Guanhua, et al.
Published: (2025)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Learning Dynamics of LLM Finetuning
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
by: Zhao, Beidi, et al.
Published: (2026)
by: Zhao, Beidi, et al.
Published: (2026)
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
by: Li, Zichao, et al.
Published: (2026)
by: Li, Zichao, et al.
Published: (2026)
Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
by: Behnia, Tina, et al.
Published: (2025)
by: Behnia, Tina, et al.
Published: (2025)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
by: Li, Yushu, et al.
Published: (2026)
by: Li, Yushu, et al.
Published: (2026)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
by: Yao, Xincheng, et al.
Published: (2026)
by: Yao, Xincheng, et al.
Published: (2026)
Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
by: Hu, Ziyou, et al.
Published: (2025)
by: Hu, Ziyou, et al.
Published: (2025)
Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
by: Xing, Shangyu, et al.
Published: (2025)
by: Xing, Shangyu, et al.
Published: (2025)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations
by: Zhao, Yize, et al.
Published: (2024)
by: Zhao, Yize, et al.
Published: (2024)
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
by: Yang, Running, et al.
Published: (2025)
by: Yang, Running, et al.
Published: (2025)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
by: Deora, Puneesh, et al.
Published: (2025)
by: Deora, Puneesh, et al.
Published: (2025)
Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
by: Cui, Wendi, et al.
Published: (2024)
by: Cui, Wendi, et al.
Published: (2024)
On the Hidden Objective Biases of Group-based Reinforcement Learning
by: Fontana, Aleksandar, et al.
Published: (2026)
by: Fontana, Aleksandar, et al.
Published: (2026)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
by: Vakilian, Vala, et al.
Published: (2025)
by: Vakilian, Vala, et al.
Published: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
by: Huo, Yifu, et al.
Published: (2026)
by: Huo, Yifu, et al.
Published: (2026)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards
by: Zhang, Jiajie, et al.
Published: (2026)
by: Zhang, Jiajie, et al.
Published: (2026)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
by: Mahdavi, Sadegh, et al.
Published: (2025)
by: Mahdavi, Sadegh, et al.
Published: (2025)
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
by: Yoon, Eunseop, et al.
Published: (2024)
by: Yoon, Eunseop, et al.
Published: (2024)
ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection
by: Gao, Changjiang, et al.
Published: (2026)
by: Gao, Changjiang, et al.
Published: (2026)
GRRM: Group Relative Reward Modeling for Machine Translation
by: Yang, Sen, et al.
Published: (2026)
by: Yang, Sen, et al.
Published: (2026)
Subliminal Steering: Stronger Encoding of Hidden Signals
by: Morgulis, George, et al.
Published: (2026)
by: Morgulis, George, et al.
Published: (2026)
Similar Items
-
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
by: Deng, Wenlong, et al.
Published: (2025) -
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
by: Deng, Wenlong, et al.
Published: (2025) -
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026) -
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
by: Deng, Wenlong, et al.
Published: (2025) -
Unlocking the Potential of Prompt-Tuning in Bridging Generalized and Personalized Federated Learning
by: Deng, Wenlong, et al.
Published: (2023)