Policy Optimization with Smooth Guidance Learned from State-Only Demonstrations
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Guojian, Wu, Faguo, Zhang, Xiao, Chen, Tianyuan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Diverse Policies with Soft Self-Generated Guidance
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Trajectory-Oriented Policy Optimization with Sparse Rewards
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Preference-Guided Reinforcement Learning for Efficient Exploration
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
by: Wang, Guojian, et al.
Published: (2023)
by: Wang, Guojian, et al.
Published: (2023)
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
by: Yao, Qingmao, et al.
Published: (2025)
by: Yao, Qingmao, et al.
Published: (2025)
Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few Demonstrations
by: Zhu, Yujie, et al.
Published: (2025)
by: Zhu, Yujie, et al.
Published: (2025)
Policy Constraint by Only Support Constraint for Offline Reinforcement Learning
by: Gao, Yunkai, et al.
Published: (2025)
by: Gao, Yunkai, et al.
Published: (2025)
Positive-Only Drifting Policy Optimization
by: Zhang, Qi
Published: (2026)
by: Zhang, Qi
Published: (2026)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2025)
by: Gao, Chen-Xiao, et al.
Published: (2025)
Imitation Learning by State-Only Distribution Matching
by: Boborzi, Damian, et al.
Published: (2022)
by: Boborzi, Damian, et al.
Published: (2022)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding States
by: Kwak, Kyoleen, et al.
Published: (2026)
by: Kwak, Kyoleen, et al.
Published: (2026)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
by: Zhang, Qingru, et al.
Published: (2025)
by: Zhang, Qingru, et al.
Published: (2025)
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
by: Chen, Xuyang, et al.
Published: (2025)
by: Chen, Xuyang, et al.
Published: (2025)
RePO: Bridging On-Policy Learning and Off-Policy Knowledge through Rephrasing Policy Optimization
by: Xia, Linxuan, et al.
Published: (2026)
by: Xia, Linxuan, et al.
Published: (2026)
Learning to Reason under Off-Policy Guidance
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
DADP: Domain Adaptive Diffusion Policy
by: Wang, Pengcheng, et al.
Published: (2026)
by: Wang, Pengcheng, et al.
Published: (2026)
Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings
by: Chen, Wenxin, et al.
Published: (2026)
by: Chen, Wenxin, et al.
Published: (2026)
Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
by: Chen, Lisha, et al.
Published: (2025)
by: Chen, Lisha, et al.
Published: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
by: Wang, Zeyuan, et al.
Published: (2025)
by: Wang, Zeyuan, et al.
Published: (2025)
Gradient Guidance for Diffusion Models: An Optimization Perspective
by: Guo, Yingqing, et al.
Published: (2024)
by: Guo, Yingqing, et al.
Published: (2024)
State-wise Constrained Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
by: Wei, Zeming, et al.
Published: (2023)
by: Wei, Zeming, et al.
Published: (2023)
Goal-Reaching Policy Learning from Non-Expert Observations via Effective Subgoal Guidance
by: Huang, RenMing, et al.
Published: (2024)
by: Huang, RenMing, et al.
Published: (2024)
REG: Rectified Gradient Guidance for Conditional Diffusion Models
by: Gao, Zhengqi, et al.
Published: (2025)
by: Gao, Zhengqi, et al.
Published: (2025)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
by: Xiao, Wei, et al.
Published: (2025)
by: Xiao, Wei, et al.
Published: (2025)
Imitation Learning from Purified Demonstrations
by: Wang, Yunke, et al.
Published: (2023)
by: Wang, Yunke, et al.
Published: (2023)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
by: Zhang, Tianqi, et al.
Published: (2025)
by: Zhang, Tianqi, et al.
Published: (2025)
Reinforcement Learning with Curriculum-inspired Adaptive Direct Policy Guidance for Truck Dispatching
by: Meng, Shi, et al.
Published: (2025)
by: Meng, Shi, et al.
Published: (2025)
Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Effect of Optimizer, Initializer, and Architecture of Hypernetworks on Continual Learning from Demonstration
by: Auddy, Sayantan, et al.
Published: (2023)
by: Auddy, Sayantan, et al.
Published: (2023)
Functional Autoencoder for Smoothing and Representation Learning
by: Wu, Sidi, et al.
Published: (2024)
by: Wu, Sidi, et al.
Published: (2024)
Data-Efficient RLVR via Off-Policy Influence Guidance
by: Zhu, Erle, et al.
Published: (2025)
by: Zhu, Erle, et al.
Published: (2025)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
by: Zhang, Jing, et al.
Published: (2023)
by: Zhang, Jing, et al.
Published: (2023)
Data Fusion-Enhanced Decision Transformer for Stable Cross-Domain Generalization
by: Wang, Guojian, et al.
Published: (2025)
by: Wang, Guojian, et al.
Published: (2025)
Modular Diffusion Policy Training: Decoupling and Recombining Guidance and Diffusion for Offline RL
by: Chen, Zhaoyang, et al.
Published: (2025)
by: Chen, Zhaoyang, et al.
Published: (2025)
Dual Approximation Policy Optimization
by: Xiong, Zhihan, et al.
Published: (2024)
by: Xiong, Zhihan, et al.
Published: (2024)
Similar Items
-
Learning Diverse Policies with Soft Self-Generated Guidance
by: Wang, Guojian, et al.
Published: (2024) -
Trajectory-Oriented Policy Optimization with Sparse Rewards
by: Wang, Guojian, et al.
Published: (2024) -
Preference-Guided Reinforcement Learning for Efficient Exploration
by: Wang, Guojian, et al.
Published: (2024) -
Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
by: Wang, Guojian, et al.
Published: (2023) -
Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood
by: Yao, Qingmao, et al.
Published: (2025)