Generalized Advantage Estimation for Distributional Policy Gradients
Fuente:
arXiv
Saved in:
| Main Authors: | Shaik, Shahil, Smereka, Jonathon M., Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent Deep Reinforcement Learning Under Constrained Communications
by: Shaik, Shahil, et al.
Published: (2026)
by: Shaik, Shahil, et al.
Published: (2026)
MA-VLCM: A Vision Language Critic Model for Value Estimation of Policies in Multi-Agent Team Settings
by: Shaik, Shahil, et al.
Published: (2026)
by: Shaik, Shahil, et al.
Published: (2026)
Actor-Critic Cooperative Compensation to Model Predictive Control for Off-Road Autonomous Vehicles Under Unknown Dynamics
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Flow Matching Policy Gradients
by: McAllister, David, et al.
Published: (2025)
by: McAllister, David, et al.
Published: (2025)
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025)
by: Liu, Xinjie, et al.
Published: (2025)
Reinforcement Learning Compensated Model Predictive Control for Off-road Driving on Unknown Deformable Terrain
by: Gupta, Prakhar, et al.
Published: (2024)
by: Gupta, Prakhar, et al.
Published: (2024)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
by: Xue, Rong, et al.
Published: (2025)
by: Xue, Rong, et al.
Published: (2025)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
by: Koo, Juil, et al.
Published: (2026)
by: Koo, Juil, et al.
Published: (2026)
Score and Distribution Matching Policy: Advanced Accelerated Visuomotor Policies via Matched Distillation
by: Jia, Bofang, et al.
Published: (2024)
by: Jia, Bofang, et al.
Published: (2024)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Latent-Conditioned Policy Gradient for Multi-Objective Deep Reinforcement Learning
by: Kanazawa, Takuya, et al.
Published: (2023)
by: Kanazawa, Takuya, et al.
Published: (2023)
Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC
by: Zhang, Xinglong, et al.
Published: (2024)
by: Zhang, Xinglong, et al.
Published: (2024)
ETGL-DDPG: A Deep Deterministic Policy Gradient Algorithm for Sparse Reward Continuous Control
by: Futuhi, Ehsan, et al.
Published: (2024)
by: Futuhi, Ehsan, et al.
Published: (2024)
Distributed Policy Gradient for Linear Quadratic Networked Control with Limited Communication Range
by: Yan, Yuzi, et al.
Published: (2024)
by: Yan, Yuzi, et al.
Published: (2024)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
by: Atreya, Pranav, et al.
Published: (2025)
by: Atreya, Pranav, et al.
Published: (2025)
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
by: Cao, Jiahang, et al.
Published: (2025)
by: Cao, Jiahang, et al.
Published: (2025)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024)
by: Xu, Kechun, et al.
Published: (2024)
Collision Probability Distribution Estimation via Temporal Difference Learning
by: Steinecker, Thomas, et al.
Published: (2024)
by: Steinecker, Thomas, et al.
Published: (2024)
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
by: Jain, Ayush, et al.
Published: (2024)
by: Jain, Ayush, et al.
Published: (2024)
Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation
by: Danesh, Mohamad H., et al.
Published: (2025)
by: Danesh, Mohamad H., et al.
Published: (2025)
Real-Time Generative Policy via Langevin-Guided Flow Matching for Autonomous Driving
by: Zhu, Tianze, et al.
Published: (2026)
by: Zhu, Tianze, et al.
Published: (2026)
Optimisation of Structured Neural Controller Based on Continuous-Time Policy Gradient
by: Cho, Namhoon, et al.
Published: (2022)
by: Cho, Namhoon, et al.
Published: (2022)
Data-efficient, Explainable and Safe Box Manipulation: Illustrating the Advantages of Physical Priors in Model-Predictive Control
by: Salehi, Achkan, et al.
Published: (2023)
by: Salehi, Achkan, et al.
Published: (2023)
Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
by: Longhini, Alberta, et al.
Published: (2026)
by: Longhini, Alberta, et al.
Published: (2026)
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
by: Patil, Sarvesh, et al.
Published: (2026)
by: Patil, Sarvesh, et al.
Published: (2026)
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
by: Schoepp, Sheila, et al.
Published: (2024)
by: Schoepp, Sheila, et al.
Published: (2024)
Solving Reach-Avoid-Stay Problems Using Deep Deterministic Policy Gradients
by: Chenevert, Gabriel, et al.
Published: (2024)
by: Chenevert, Gabriel, et al.
Published: (2024)
Equivariant Diffusion Policy
by: Wang, Dian, et al.
Published: (2024)
by: Wang, Dian, et al.
Published: (2024)
Foundational Policy Acquisition via Multitask Learning for Motor Skill Generation
by: Yamamori, Satoshi, et al.
Published: (2023)
by: Yamamori, Satoshi, et al.
Published: (2023)
Diffusion Policy Policy Optimization
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
DADP: Domain Adaptive Diffusion Policy
by: Wang, Pengcheng, et al.
Published: (2026)
by: Wang, Pengcheng, et al.
Published: (2026)
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
by: Huang, Haojie, et al.
Published: (2024)
by: Huang, Haojie, et al.
Published: (2024)
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
by: Odonchimed, Sodtavilan, et al.
Published: (2025)
by: Odonchimed, Sodtavilan, et al.
Published: (2025)
FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation
by: Welte, Edgar, et al.
Published: (2026)
by: Welte, Edgar, et al.
Published: (2026)
Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments
by: Etukuru, Haritheja, et al.
Published: (2024)
by: Etukuru, Haritheja, et al.
Published: (2024)
Curriculum Imitation Learning of Distributed Multi-Robot Policies
by: Roche, Jesús, et al.
Published: (2025)
by: Roche, Jesús, et al.
Published: (2025)
Similar Items
-
Multi-Agent Deep Reinforcement Learning Under Constrained Communications
by: Shaik, Shahil, et al.
Published: (2026) -
MA-VLCM: A Vision Language Critic Model for Value Estimation of Policies in Multi-Agent Team Settings
by: Shaik, Shahil, et al.
Published: (2026) -
Actor-Critic Cooperative Compensation to Model Predictive Control for Off-Road Autonomous Vehicles Under Unknown Dynamics
by: Gupta, Prakhar, et al.
Published: (2025) -
Flow Matching Policy Gradients
by: McAllister, David, et al.
Published: (2025) -
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025)