Trajectory-Oriented Policy Optimization with Sparse Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Guojian, Wu, Faguo, Zhang, Xiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
von: Wang, Guojian, et al.
Veröffentlicht: (2023)
Preference-Guided Reinforcement Learning for Efficient Exploration
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
von: Wang, Guojian, et al.
Veröffentlicht: (2024)
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
von: Cho, Minjae, et al.
Veröffentlicht: (2026)
von: Cho, Minjae, et al.
Veröffentlicht: (2026)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
von: Sana, Mohamed, et al.
Veröffentlicht: (2026)
TMS: Trajectory-Mixed Supervision for Reward-Free, On-Policy SFT
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2026)
von: Khan, Rana Muhammad Shahroz, et al.
Veröffentlicht: (2026)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
Interpretable Reward Model via Sparse Autoencoder
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shuyi, et al.
Veröffentlicht: (2025)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
von: Bai, Yang, et al.
Veröffentlicht: (2026)
von: Bai, Yang, et al.
Veröffentlicht: (2026)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
GOPO: Policy Optimization using Ranked Rewards
von: Choi, Kyuseong, et al.
Veröffentlicht: (2026)
von: Choi, Kyuseong, et al.
Veröffentlicht: (2026)
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
von: Xu, Wenyuan, et al.
Veröffentlicht: (2025)
von: Xu, Wenyuan, et al.
Veröffentlicht: (2025)
Learning Fill-in Reduction Ordering via Graph Policy Optimization for Sparse Matrices
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
DADP: Domain Adaptive Diffusion Policy
von: Wang, Pengcheng, et al.
Veröffentlicht: (2026)
von: Wang, Pengcheng, et al.
Veröffentlicht: (2026)
Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoying, et al.
Veröffentlicht: (2024)
Intra-Trajectory Consistency for Reward Modeling
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
DRL-Enabled Trajectory Planing for UAV-Assisted VLC: Optimal Altitude and Reward Design
von: Lin, Tian-Tian, et al.
Veröffentlicht: (2026)
von: Lin, Tian-Tian, et al.
Veröffentlicht: (2026)
Co-Evolution of Policy and Internal Reward for Language Agents
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
von: Wang, Yixiao, et al.
Veröffentlicht: (2024)
von: Wang, Yixiao, et al.
Veröffentlicht: (2024)
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings
von: Powell, Keenan, et al.
Veröffentlicht: (2026)
von: Powell, Keenan, et al.
Veröffentlicht: (2026)
ROAD: Responsibility-Oriented Reward Design for Reinforcement Learning in Autonomous Driving
von: Chen, Yongming, et al.
Veröffentlicht: (2025)
von: Chen, Yongming, et al.
Veröffentlicht: (2025)
Value-Free Policy Optimization via Reward Partitioning
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
von: Faye, Bilal, et al.
Veröffentlicht: (2025)
ETGL-DDPG: A Deep Deterministic Policy Gradient Algorithm for Sparse Reward Continuous Control
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2024)
von: Futuhi, Ehsan, et al.
Veröffentlicht: (2024)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
Policy Learning for Balancing Short-Term and Long-Term Rewards
von: Wu, Peng, et al.
Veröffentlicht: (2024)
von: Wu, Peng, et al.
Veröffentlicht: (2024)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
Attention-Based Reward Shaping for Sparse and Delayed Rewards
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
von: Holmes, Ian, et al.
Veröffentlicht: (2025)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
von: Matrenok, Simon, et al.
Veröffentlicht: (2025)
von: Matrenok, Simon, et al.
Veröffentlicht: (2025)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2024)
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2024)
Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias
von: Hara, Rahaf Abu, et al.
Veröffentlicht: (2026)
von: Hara, Rahaf Abu, et al.
Veröffentlicht: (2026)
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zheng, et al.
Veröffentlicht: (2026)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
von: Wang, Guojian, et al.
Veröffentlicht: (2023) -
Learning Diverse Policies with Soft Self-Generated Guidance
von: Wang, Guojian, et al.
Veröffentlicht: (2024) -
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
von: Wang, Guojian, et al.
Veröffentlicht: (2023) -
Preference-Guided Reinforcement Learning for Efficient Exploration
von: Wang, Guojian, et al.
Veröffentlicht: (2024) -
Intrinsic Reward Policy Optimization for Sparse-Reward Environments
von: Cho, Minjae, et al.
Veröffentlicht: (2026)