Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Ayush, Kosaka, Norio, Li, Xinhu, Kim, Kyung-Min, Bıyık, Erdem, Lim, Joseph J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025)
by: Li, Xinhu, et al.
Published: (2025)
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing
by: Zhang, Grace, et al.
Published: (2023)
by: Zhang, Grace, et al.
Published: (2023)
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025)
by: Korkmaz, Yigit, et al.
Published: (2025)
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)
by: Banayeeanzade, Amin, et al.
Published: (2025)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Multi-Agent Inverse Q-Learning from Demonstrations
by: Haynam, Nathaniel, et al.
Published: (2025)
by: Haynam, Nathaniel, et al.
Published: (2025)
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
by: Hwang, Minjune, et al.
Published: (2026)
by: Hwang, Minjune, et al.
Published: (2026)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)
by: Zhang, Jesse, et al.
Published: (2025)
IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models
by: Ling, Yiyang, et al.
Published: (2025)
by: Ling, Yiyang, et al.
Published: (2025)
CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations
by: Liang, Anthony, et al.
Published: (2025)
by: Liang, Anthony, et al.
Published: (2025)
A Generalized Acquisition Function for Preference-based Reward Learning
by: Ellis, Evan, et al.
Published: (2024)
by: Ellis, Evan, et al.
Published: (2024)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
Robust Policy Learning via Offline Skill Diffusion
by: Kim, Woo Kyung, et al.
Published: (2024)
by: Kim, Woo Kyung, et al.
Published: (2024)
In-Context Policy Adaptation via Cross-Domain Skill Diffusion
by: Yoo, Minjong, et al.
Published: (2025)
by: Yoo, Minjong, et al.
Published: (2025)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
by: Koo, Juil, et al.
Published: (2026)
by: Koo, Juil, et al.
Published: (2026)
Reward Learning from Suboptimal Demonstrations with Applications in Surgical Electrocautery
by: Karimi, Zohre, et al.
Published: (2024)
by: Karimi, Zohre, et al.
Published: (2024)
SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling
by: Zhang, Jesse, et al.
Published: (2023)
by: Zhang, Jesse, et al.
Published: (2023)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
by: Yang, Brian, et al.
Published: (2024)
by: Yang, Brian, et al.
Published: (2024)
Attention-Based Neural-Augmented Kalman Filter for Legged Robot State Estimation
by: Lee, Seokju, et al.
Published: (2026)
by: Lee, Seokju, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
Trust Region Q Adjoint Matching
by: Dong, Yonghoon, et al.
Published: (2026)
by: Dong, Yonghoon, et al.
Published: (2026)
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning
by: Alles, Marvin, et al.
Published: (2025)
by: Alles, Marvin, et al.
Published: (2025)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
ViSaRL: Visual Reinforcement Learning Guided by Human Saliency
by: Liang, Anthony, et al.
Published: (2024)
by: Liang, Anthony, et al.
Published: (2024)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025)
by: Liu, Xinjie, et al.
Published: (2025)
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
by: Schoepp, Sheila, et al.
Published: (2024)
by: Schoepp, Sheila, et al.
Published: (2024)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
by: Cai, Shizhe, et al.
Published: (2025)
by: Cai, Shizhe, et al.
Published: (2025)
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
by: Liang, Anthony, et al.
Published: (2026)
by: Liang, Anthony, et al.
Published: (2026)
Distilling Reinforcement Learning Policies for Interpretable Robot Locomotion: Gradient Boosting Machines and Symbolic Regression
by: Acero, Fernando, et al.
Published: (2024)
by: Acero, Fernando, et al.
Published: (2024)
RAILGUN: A Unified Convolutional Policy for Multi-Agent Path Finding Across Different Environments and Tasks
by: Tang, Yimin, et al.
Published: (2025)
by: Tang, Yimin, et al.
Published: (2025)
Don't Start from Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion
by: Chen, Kaiqi, et al.
Published: (2024)
by: Chen, Kaiqi, et al.
Published: (2024)
Grasp Anything: Combining Teacher-Augmented Policy Gradient Learning with Instance Segmentation to Grasp Arbitrary Objects
by: Mosbach, Malte, et al.
Published: (2024)
by: Mosbach, Malte, et al.
Published: (2024)
Multi-Agent Reinforcement Learning for Unmanned Aerial Vehicle Coordination by Multi-Critic Policy Gradient Optimization
by: Alon, Yoav, et al.
Published: (2020)
by: Alon, Yoav, et al.
Published: (2020)
Decoupled Q-Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Q-learning with Adjoint Matching
by: Li, Qiyang, et al.
Published: (2026)
by: Li, Qiyang, et al.
Published: (2026)
Similar Items
-
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
by: Li, Xinhu, et al.
Published: (2025) -
Actor-Free Continuous Control via Structurally Maximizable Q-Functions
by: Korkmaz, Yigit, et al.
Published: (2025) -
QMP: Q-switch Mixture of Policies for Multi-Task Behavior Sharing
by: Zhang, Grace, et al.
Published: (2023) -
MILE: Model-based Intervention Learning
by: Korkmaz, Yigit, et al.
Published: (2025) -
GABRIL: Gaze-Based Regularization for Mitigating Causal Confusion in Imitation Learning
by: Banayeeanzade, Amin, et al.
Published: (2025)