Intrinsic Reward Policy Optimization for Sparse-Reward Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Minjae, Tran, Huy Trong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contraction Actor-Critic: Contraction Metric-Guided Reinforcement Learning for Robust Path Tracking
von: Cho, Minjae, et al.
Veröffentlicht: (2025)
von: Cho, Minjae, et al.
Veröffentlicht: (2025)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
von: Lyu, Guwen, et al.
Veröffentlicht: (2024)
von: Lyu, Guwen, et al.
Veröffentlicht: (2024)
GOPO: Policy Optimization using Ranked Rewards
von: Choi, Kyuseong, et al.
Veröffentlicht: (2026)
von: Choi, Kyuseong, et al.
Veröffentlicht: (2026)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
Accelerating Proximal Policy Optimization Learning Using Task Prediction for Solving Environments with Delayed Rewards
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
von: Ahmad, Ahmad, et al.
Veröffentlicht: (2024)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
IRIS: Intrinsic Reward Image Synthesis
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
von: Chen, Yihang, et al.
Veröffentlicht: (2025)
What Fundamental Structure in Reward Functions Enables Efficient Sparse-Reward Learning?
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning
von: Nguyen, Viet Bac, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet Bac, et al.
Veröffentlicht: (2026)
Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization
von: Wu, Boyuan
Veröffentlicht: (2025)
von: Wu, Boyuan
Veröffentlicht: (2025)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
Fairness Aware Reward Optimization
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2026)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
von: Wen, Xuexiang, et al.
Veröffentlicht: (2026)
BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
von: Lidayan, Aly, et al.
Veröffentlicht: (2024)
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
WARP: On the Benefits of Weight Averaged Rewarded Policies
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
Entropy Centroids as Intrinsic Rewards for Test-Time Scaling
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
von: Zhao, Wenshuo, et al.
Veröffentlicht: (2026)
ReDit: Reward Dithering for Improved LLM Policy Optimization
von: Wei, Chenxing, et al.
Veröffentlicht: (2025)
von: Wei, Chenxing, et al.
Veröffentlicht: (2025)
2048: Reinforcement Learning in a Delayed Reward Environment
von: Saligram, Prady, et al.
Veröffentlicht: (2025)
von: Saligram, Prady, et al.
Veröffentlicht: (2025)
Hierarchical Meta-Reinforcement Learning via Automated Macro-Action Discovery
von: Cho, Minjae, et al.
Veröffentlicht: (2024)
von: Cho, Minjae, et al.
Veröffentlicht: (2024)
Sparsity-based Safety Conservatism for Constrained Offline Reinforcement Learning
von: Cho, Minjae, et al.
Veröffentlicht: (2024)
von: Cho, Minjae, et al.
Veröffentlicht: (2024)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
von: Yunis, David, et al.
Veröffentlicht: (2023)
von: Yunis, David, et al.
Veröffentlicht: (2023)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
von: Diaz-Bone, Leander, et al.
Veröffentlicht: (2025)
von: Diaz-Bone, Leander, et al.
Veröffentlicht: (2025)
Mutual-Taught for Co-adapting Policy and Reward Models
von: Shi, Tianyuan, et al.
Veröffentlicht: (2025)
von: Shi, Tianyuan, et al.
Veröffentlicht: (2025)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
von: Gumbsch, Christian, et al.
Veröffentlicht: (2026)
von: Gumbsch, Christian, et al.
Veröffentlicht: (2026)
$i$REPO: $i$mplicit Reward Pairwise Difference based Empirical Preference Optimization
von: Le, Long Tan, et al.
Veröffentlicht: (2024)
von: Le, Long Tan, et al.
Veröffentlicht: (2024)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
von: Mahrooghi, Ilia, et al.
Veröffentlicht: (2026)
von: Mahrooghi, Ilia, et al.
Veröffentlicht: (2026)
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
von: Li, Wenyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Contraction Actor-Critic: Contraction Metric-Guided Reinforcement Learning for Robust Path Tracking
von: Cho, Minjae, et al.
Veröffentlicht: (2025) -
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
von: Sana, Mohamed, et al.
Veröffentlicht: (2026) -
Attention-Based Reward Shaping for Sparse and Delayed Rewards
von: Holmes, Ian, et al.
Veröffentlicht: (2025) -
Generalized Back-Stepping Experience Replay in Sparse-Reward Environments
von: Lyu, Guwen, et al.
Veröffentlicht: (2024) -
GOPO: Policy Optimization using Ranked Rewards
von: Choi, Kyuseong, et al.
Veröffentlicht: (2026)