Improving Value Estimation Critically Enhances Vanilla Policy Gradient
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Tao, Zhang, Ruipeng, Gao, Sicun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
by: Li, Xidong, et al.
Published: (2026)
by: Li, Xidong, et al.
Published: (2026)
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
SPACeR: Self-Play Anchoring with Centralized Reference Models
by: Chang, Wei-Jer, et al.
Published: (2025)
by: Chang, Wei-Jer, et al.
Published: (2025)
CUPID: Curating Data your Robot Loves with Influence Functions
by: Agia, Christopher, et al.
Published: (2025)
by: Agia, Christopher, et al.
Published: (2025)
SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action
by: Shi, Xiang, et al.
Published: (2026)
by: Shi, Xiang, et al.
Published: (2026)
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
by: Haklidir, Mehmet
Published: (2026)
by: Haklidir, Mehmet
Published: (2026)
Deep Reinforcement Learning for Adverse Garage Scenario Generation
by: Li, Kai
Published: (2024)
by: Li, Kai
Published: (2024)
Towards a Robust Soft Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms
by: Alhakami, Mohannad, et al.
Published: (2024)
by: Alhakami, Mohannad, et al.
Published: (2024)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
by: Dai, Yanning, et al.
Published: (2026)
by: Dai, Yanning, et al.
Published: (2026)
Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
by: Wang, Zizhao, et al.
Published: (2024)
by: Wang, Zizhao, et al.
Published: (2024)
Failure Prediction at Runtime for Generative Robot Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Adversarial Constrained Policy Optimization: Improving Constrained Reinforcement Learning by Adapting Budgets
by: Ma, Jianmina, et al.
Published: (2024)
by: Ma, Jianmina, et al.
Published: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
Positive-Only Drifting Policy Optimization
by: Zhang, Qi
Published: (2026)
by: Zhang, Qi
Published: (2026)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
by: Agia, Christopher, et al.
Published: (2024)
by: Agia, Christopher, et al.
Published: (2024)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
by: Ai, Bo, et al.
Published: (2024)
by: Ai, Bo, et al.
Published: (2024)
DSSE: a drone swarm search environment
by: Castanares, Manuel, et al.
Published: (2023)
by: Castanares, Manuel, et al.
Published: (2023)
Hierarchical Universal Value Function Approximators
by: Arora, Rushiv
Published: (2024)
by: Arora, Rushiv
Published: (2024)
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
The Reality Gap in Robotics: Challenges, Solutions, and Best Practices
by: Aljalbout, Elie, et al.
Published: (2025)
by: Aljalbout, Elie, et al.
Published: (2025)
Achieving Scalable Robot Autonomy via neurosymbolic planning using lightweight local LLM
by: Attolino, Nicholas, et al.
Published: (2025)
by: Attolino, Nicholas, et al.
Published: (2025)
The Shortcomings of Force-from-Motion in Robot Learning
by: Aljalbout, Elie, et al.
Published: (2024)
by: Aljalbout, Elie, et al.
Published: (2024)
A Framework for Neurosymbolic Robot Action Planning using Large Language Models
by: Capitanelli, Alessio, et al.
Published: (2023)
by: Capitanelli, Alessio, et al.
Published: (2023)
The State of Robot Motion Generation
by: Bekris, Kostas E., et al.
Published: (2024)
by: Bekris, Kostas E., et al.
Published: (2024)
Cost-Aware Diffusion Active Search
by: Banerjee, Arundhati, et al.
Published: (2026)
by: Banerjee, Arundhati, et al.
Published: (2026)
Physics-Informed Policy Optimization via Analytic Dynamics Regularization
by: Chandra, Namai, et al.
Published: (2026)
by: Chandra, Namai, et al.
Published: (2026)
Expressive Value Learning for Scalable Offline Reinforcement Learning
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Atomic-Probe Governance for Skill Updates in Compositional Robot Policies
by: Qin, Xue, et al.
Published: (2026)
by: Qin, Xue, et al.
Published: (2026)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
SAFE-SIM: Safety-Critical Closed-Loop Traffic Simulation with Diffusion-Controllable Adversaries
by: Chang, Wei-Jer, et al.
Published: (2023)
by: Chang, Wei-Jer, et al.
Published: (2023)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Altered Thoughts, Altered Actions: Probing Chain-of-Thought Vulnerabilities in VLA Robotic Manipulation
by: Trinh, Tuan Duong, et al.
Published: (2026)
by: Trinh, Tuan Duong, et al.
Published: (2026)
Multi-Task Reinforcement Learning with Language-Encoded Gated Policy Networks
by: Arora, Rushiv
Published: (2025)
by: Arora, Rushiv
Published: (2025)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
by: Ged, François, et al.
Published: (2023)
by: Ged, François, et al.
Published: (2023)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Pure Planning to Pure Policies and In Between with a Recursive Tree Planner
by: Redlich, A. Norman
Published: (2024)
by: Redlich, A. Norman
Published: (2024)
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Intervention-Assisted Policy Gradient Methods for Online Stochastic Queuing Network Optimization: Technical Report
by: Wigmore, Jerrod, et al.
Published: (2024)
by: Wigmore, Jerrod, et al.
Published: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration
by: Rodemann, Julian, et al.
Published: (2024)
by: Rodemann, Julian, et al.
Published: (2024)
Similar Items
-
MTDrive: Multi-turn Interactive Reinforcement Learning for Autonomous Driving
by: Li, Xidong, et al.
Published: (2026) -
MEReQ: Max-Ent Residual-Q Inverse RL for Sample-Efficient Alignment from Intervention
by: Chen, Yuxin, et al.
Published: (2024) -
SPACeR: Self-Play Anchoring with Centralized Reference Models
by: Chang, Wei-Jer, et al.
Published: (2025) -
CUPID: Curating Data your Robot Loves with Influence Functions
by: Agia, Christopher, et al.
Published: (2025) -
SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action
by: Shi, Xiang, et al.
Published: (2026)