Towards better dense rewards in Reinforcement Learning Applications
Fuente:
arXiv
Saved in:
| Main Author: | Zhang, Shuyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Episodic Reinforcement Learning with Expanded State-reward Space
by: Liang, Dayang, et al.
Published: (2024)
by: Liang, Dayang, et al.
Published: (2024)
The impact of intrinsic rewards on exploration in Reinforcement Learning
by: Kayal, Aya, et al.
Published: (2025)
by: Kayal, Aya, et al.
Published: (2025)
Deep Reinforcement Learning with anticipatory reward in LSTM for Collision Avoidance of Mobile Robots
by: Poulet, Olivier, et al.
Published: (2025)
by: Poulet, Olivier, et al.
Published: (2025)
Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
by: Zhang, Shuyuan, et al.
Published: (2025)
by: Zhang, Shuyuan, et al.
Published: (2025)
Self-rewarding correction for mathematical reasoning
by: Xiong, Wei, et al.
Published: (2025)
by: Xiong, Wei, et al.
Published: (2025)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
by: Arabzadeh, Negar, et al.
Published: (2024)
by: Arabzadeh, Negar, et al.
Published: (2024)
EVAL: EigenVector-based Average-reward Learning
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
What should be observed for optimal reward in POMDPs?
by: Konsta, Alyzia-Maria, et al.
Published: (2024)
by: Konsta, Alyzia-Maria, et al.
Published: (2024)
Streaming Looking Ahead with Token-level Self-reward
by: Zhang, Hongming, et al.
Published: (2025)
by: Zhang, Hongming, et al.
Published: (2025)
Noise-based reward-modulated learning
by: Fernández, Jesús García, et al.
Published: (2025)
by: Fernández, Jesús García, et al.
Published: (2025)
Active teacher selection for reward learning
by: Freedman, Rachel, et al.
Published: (2023)
by: Freedman, Rachel, et al.
Published: (2023)
MOSLIM:Align with diverse preferences in prompts through reward classification
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Information-theoretic analysis of world models in optimal reward maximizers
by: Harwood, Alfred, et al.
Published: (2026)
by: Harwood, Alfred, et al.
Published: (2026)
Reinforcement Learning with Partial Parametric Model Knowledge
by: Wang, Shuyuan, et al.
Published: (2023)
by: Wang, Shuyuan, et al.
Published: (2023)
Zero Reinforcement Learning Towards General Domains
by: Zeng, Yuyuan, et al.
Published: (2025)
by: Zeng, Yuyuan, et al.
Published: (2025)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
by: Yan, Xue, et al.
Published: (2023)
by: Yan, Xue, et al.
Published: (2023)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
Toward Virtuous Reinforcement Learning: A Critique and Roadmap
by: Ghasemi, Majid, et al.
Published: (2025)
by: Ghasemi, Majid, et al.
Published: (2025)
R-ParVI: Particle-based variational inference through lens of rewards
by: Huang, Yongchao
Published: (2025)
by: Huang, Yongchao
Published: (2025)
Self-supervised network distillation: an effective approach to exploration in sparse reward environments
by: Pecháč, Matej, et al.
Published: (2023)
by: Pecháč, Matej, et al.
Published: (2023)
Towards General-Purpose Model-Free Reinforcement Learning
by: Fujimoto, Scott, et al.
Published: (2025)
by: Fujimoto, Scott, et al.
Published: (2025)
Towards Robust Offline Reinforcement Learning under Diverse Data Corruption
by: Yang, Rui, et al.
Published: (2023)
by: Yang, Rui, et al.
Published: (2023)
Towards Interpretable Deep Reinforcement Learning Models via Inverse Reinforcement Learning
by: Xie, Sean, et al.
Published: (2022)
by: Xie, Sean, et al.
Published: (2022)
Unifying Causal Reinforcement Learning: Survey, Taxonomy, Algorithms and Applications
by: Cunha, Cristiano da Costa, et al.
Published: (2025)
by: Cunha, Cristiano da Costa, et al.
Published: (2025)
Self-Driving Car Racing: Application of Deep Reinforcement Learning
by: Yuwono, Florentiana, et al.
Published: (2024)
by: Yuwono, Florentiana, et al.
Published: (2024)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
reward-lens: A Mechanistic Interpretability Library for Reward Models
by: Nadaf, Mohammed Suhail B
Published: (2026)
by: Nadaf, Mohammed Suhail B
Published: (2026)
A Survey on Applications of Reinforcement Learning in Spatial Resource Allocation
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
by: Wang, Haiming, et al.
Published: (2025)
by: Wang, Haiming, et al.
Published: (2025)
Towards User-level Private Reinforcement Learning with Human Feedback
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Federated Reinforcement Learning for Runtime Optimization of AI Applications in Smart Eyewears
by: Sedghani, Hamta, et al.
Published: (2025)
by: Sedghani, Hamta, et al.
Published: (2025)
Deep Reinforcement Learning Based Systems for Safety Critical Applications in Aerospace
by: Sherifi, Abedin
Published: (2024)
by: Sherifi, Abedin
Published: (2024)
Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges
by: Zhou, Ziyuan, et al.
Published: (2023)
by: Zhou, Ziyuan, et al.
Published: (2023)
A Review of Reinforcement Learning in Financial Applications
by: Bai, Yahui, et al.
Published: (2024)
by: Bai, Yahui, et al.
Published: (2024)
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Content Bias in Deep Learning Image Age Approximation: A new Approach Towards better Explainability
by: Jöchl, Robert, et al.
Published: (2023)
by: Jöchl, Robert, et al.
Published: (2023)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
SCAR: Shapley Credit Assignment for More Efficient RLHF
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Similar Items
-
Episodic Reinforcement Learning with Expanded State-reward Space
by: Liang, Dayang, et al.
Published: (2024) -
The impact of intrinsic rewards on exploration in Reinforcement Learning
by: Kayal, Aya, et al.
Published: (2025) -
Deep Reinforcement Learning with anticipatory reward in LSTM for Collision Avoidance of Mobile Robots
by: Poulet, Olivier, et al.
Published: (2025) -
Incorporating Spatial Information into Goal-Conditioned Hierarchical Reinforcement Learning via Graph Representations
by: Zhang, Shuyuan, et al.
Published: (2025) -
Self-rewarding correction for mathematical reasoning
by: Xiong, Wei, et al.
Published: (2025)