Where-to-Learn: Analytical Policy Gradient Directed Exploration for On-Policy Robotic Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Leixin, Yao, Xinchen, Liu, Ben, Yang, Liangjing, Chen, Hua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Robustness: Learning Unknown Dynamic Load Adaptation for Quadruped Locomotion on Rough Terrain
by: Chang, Leixin, et al.
Published: (2025)
by: Chang, Leixin, et al.
Published: (2025)
PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning
by: Yang, Shunpeng, et al.
Published: (2026)
by: Yang, Shunpeng, et al.
Published: (2026)
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Diffusion Policy through Conditional Proximal Policy Optimization
by: Liu, Ben, et al.
Published: (2026)
by: Liu, Ben, et al.
Published: (2026)
Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
by: Nie, Buqing, et al.
Published: (2025)
by: Nie, Buqing, et al.
Published: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
Learning Agile Gate Traversal via Analytical Optimal Policy Gradient
by: Sun, Tianchen, et al.
Published: (2025)
by: Sun, Tianchen, et al.
Published: (2025)
Can Tabular Foundation Models Guide Exploration in Robot Policy Learning?
by: Ou, Buqing, et al.
Published: (2026)
by: Ou, Buqing, et al.
Published: (2026)
Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
by: Hu, Kaizhe, et al.
Published: (2025)
by: Hu, Kaizhe, et al.
Published: (2025)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
by: Acero, Fernando, et al.
Published: (2024)
by: Acero, Fernando, et al.
Published: (2024)
DARE: Diffusion Policy for Autonomous Robot Exploration
by: Cao, Yuhong, et al.
Published: (2024)
by: Cao, Yuhong, et al.
Published: (2024)
Flow Policy Gradients for Robot Control
by: Yi, Brent, et al.
Published: (2026)
by: Yi, Brent, et al.
Published: (2026)
Latent-Conditioned Policy Gradient for Multi-Objective Deep Reinforcement Learning
by: Kanazawa, Takuya, et al.
Published: (2023)
by: Kanazawa, Takuya, et al.
Published: (2023)
DAPPER: Discriminability-Aware Policy-to-Policy Preference-Based Reinforcement Learning for Query-Efficient Robot Skill Acquisition
by: Kadokawa, Yuki, et al.
Published: (2025)
by: Kadokawa, Yuki, et al.
Published: (2025)
GCNT: Graph-Based Transformer Policies for Morphology-Agnostic Reinforcement Learning
by: Luo, Yingbo, et al.
Published: (2025)
by: Luo, Yingbo, et al.
Published: (2025)
Finetuning Deep Reinforcement Learning Policies with Evolutionary Strategies for Control of Underactuated Robots
by: Calì, Marco, et al.
Published: (2025)
by: Calì, Marco, et al.
Published: (2025)
Poke and Strike: Learning Task-Informed Exploration Policies
by: Aoyama, Marina Y., et al.
Published: (2025)
by: Aoyama, Marina Y., et al.
Published: (2025)
NFPO: Stabilized Policy Optimization of Normalizing Flow for Robotic Policy Learning
by: Shi, Diyuan, et al.
Published: (2026)
by: Shi, Diyuan, et al.
Published: (2026)
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
by: Zhu, Xiang, et al.
Published: (2025)
by: Zhu, Xiang, et al.
Published: (2025)
Enhancing Sample Efficiency and Exploration in Reinforcement Learning through the Integration of Diffusion Models and Proximal Policy Optimization
by: Gao, Tianci, et al.
Published: (2024)
by: Gao, Tianci, et al.
Published: (2024)
Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning
by: Shi, Diyuan, et al.
Published: (2025)
by: Shi, Diyuan, et al.
Published: (2025)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Cost-Aware Query Policies in Active Learning for Efficient Autonomous Robotic Exploration
by: Akins, Sapphira, et al.
Published: (2024)
by: Akins, Sapphira, et al.
Published: (2024)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
by: Chen, Zijia, et al.
Published: (2026)
by: Chen, Zijia, et al.
Published: (2026)
Learning Diffusion Policy from Primitive Skills for Robot Manipulation
by: Gu, Zhihao, et al.
Published: (2026)
by: Gu, Zhihao, et al.
Published: (2026)
Formal Methods in Robot Policy Learning and Verification: A Survey on Current Techniques and Future Directions
by: Manganaris, Anastasios, et al.
Published: (2026)
by: Manganaris, Anastasios, et al.
Published: (2026)
Gradient-based Regularization for Action Smoothness in Robotic Control with Reinforcement Learning
by: Lee, I, et al.
Published: (2024)
by: Lee, I, et al.
Published: (2024)
Hybrid Diffusion Policies with Projective Geometric Algebra for Efficient Robot Manipulation Learning
by: Sun, Xiatao, et al.
Published: (2025)
by: Sun, Xiatao, et al.
Published: (2025)
Enabling Stateful Behaviors for Diffusion-based Policy Learning
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
Formulating Reinforcement Learning for Human-Robot Collaboration through Off-Policy Evaluation
by: Singh, Saurav, et al.
Published: (2026)
by: Singh, Saurav, et al.
Published: (2026)
Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
by: Wang, Sen, et al.
Published: (2026)
by: Wang, Sen, et al.
Published: (2026)
Learning Robust Control Policies for Inverted Pose on Miniature Blimp Robots
by: Yang, Yuanlin, et al.
Published: (2026)
by: Yang, Yuanlin, et al.
Published: (2026)
Deep Reinforcement Learning-based Large-scale Robot Exploration
by: Cao, Yuhong, et al.
Published: (2024)
by: Cao, Yuhong, et al.
Published: (2024)
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
by: Schoepp, Sheila, et al.
Published: (2024)
by: Schoepp, Sheila, et al.
Published: (2024)
Robot Policy Transfer with Online Demonstrations: An Active Reinforcement Learning Approach
by: Hou, Muhan, et al.
Published: (2025)
by: Hou, Muhan, et al.
Published: (2025)
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
by: Jiang, Zhennan, et al.
Published: (2025)
by: Jiang, Zhennan, et al.
Published: (2025)
Learning Human-Aware Robot Policies for Adaptive Assistance
by: Qin, Jason, et al.
Published: (2024)
by: Qin, Jason, et al.
Published: (2024)
Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots
by: Liang, Junyang, et al.
Published: (2026)
by: Liang, Junyang, et al.
Published: (2026)
Similar Items
-
Beyond Robustness: Learning Unknown Dynamic Load Adaptation for Quadruped Locomotion on Rough Terrain
by: Chang, Leixin, et al.
Published: (2025) -
PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning
by: Yang, Shunpeng, et al.
Published: (2026) -
Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
by: Wang, Yi, et al.
Published: (2026) -
Diffusion Policy through Conditional Proximal Policy Optimization
by: Liu, Ben, et al.
Published: (2026) -
Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
by: Nie, Buqing, et al.
Published: (2025)