Extremum-Seeking Action Selection for Accelerating Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Ya-Chien, Gao, Sicun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025)
by: Zhang, Ruipeng, et al.
Published: (2025)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
by: Zhang, Chen Bo Calvin, et al.
Published: (2024)
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023)
by: Wang, Tao, et al.
Published: (2023)
Select before Act: Spatially Decoupled Action Repetition for Continuous Control
by: Nie, Buqing, et al.
Published: (2025)
by: Nie, Buqing, et al.
Published: (2025)
Activation-Descent Regularization for Input Optimization of ReLU Networks
by: Yu, Hongzhan, et al.
Published: (2024)
by: Yu, Hongzhan, et al.
Published: (2024)
Real-Time Execution of Action Chunking Flow Policies
by: Black, Kevin, et al.
Published: (2025)
by: Black, Kevin, et al.
Published: (2025)
Evolutionary Policy Optimization
by: Wang, Jianren, et al.
Published: (2025)
by: Wang, Jianren, et al.
Published: (2025)
Absolute Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
by: Terence, Ng Wen Zheng, et al.
Published: (2024)
by: Terence, Ng Wen Zheng, et al.
Published: (2024)
Learning Quadruped Walking from Seconds of Demonstration
by: Zhang, Ruipeng, et al.
Published: (2026)
by: Zhang, Ruipeng, et al.
Published: (2026)
Autoregressive Action Sequence Learning for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
OAT: Ordered Action Tokenization
by: Liu, Chaoqi, et al.
Published: (2026)
by: Liu, Chaoqi, et al.
Published: (2026)
Adaptive Diffusion Policy Optimization for Robotic Manipulation
by: Jiang, Huiyun, et al.
Published: (2025)
by: Jiang, Huiyun, et al.
Published: (2025)
Guided Policy Optimization under Partial Observability
by: Li, Yueheng, et al.
Published: (2025)
by: Li, Yueheng, et al.
Published: (2025)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
by: Lu, Dekun, et al.
Published: (2025)
by: Lu, Dekun, et al.
Published: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection
by: Anwar, Abrar, et al.
Published: (2025)
by: Anwar, Abrar, et al.
Published: (2025)
COSBO: Conservative Offline Simulation-Based Policy Optimization
by: Kargar, Eshagh, et al.
Published: (2024)
by: Kargar, Eshagh, et al.
Published: (2024)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
by: Guo, Jian-Ting, et al.
Published: (2025)
by: Guo, Jian-Ting, et al.
Published: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
Continual Driving Policy Optimization with Closed-Loop Individualized Curricula
by: Niu, Haoyi, et al.
Published: (2023)
by: Niu, Haoyi, et al.
Published: (2023)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
Towards Robust Policy: Enhancing Offline Reinforcement Learning with Adversarial Attacks and Defenses
by: Nguyen, Thanh, et al.
Published: (2024)
by: Nguyen, Thanh, et al.
Published: (2024)
Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics
by: Li, Chenhao, et al.
Published: (2025)
by: Li, Chenhao, et al.
Published: (2025)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
by: Bu, Qingwen, et al.
Published: (2025)
by: Bu, Qingwen, et al.
Published: (2025)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
Multi-Agent Reinforcement Learning for Unmanned Aerial Vehicle Coordination by Multi-Critic Policy Gradient Optimization
by: Alon, Yoav, et al.
Published: (2020)
by: Alon, Yoav, et al.
Published: (2020)
Behavior Generation with Latent Actions
by: Lee, Seungjae, et al.
Published: (2024)
by: Lee, Seungjae, et al.
Published: (2024)
Reinforcement Learning with Action Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
by: Ye, Weirui, et al.
Published: (2025)
by: Ye, Weirui, et al.
Published: (2025)
Unsupervised Learning of Effective Actions in Robotics
by: Zaric, Marko, et al.
Published: (2024)
by: Zaric, Marko, et al.
Published: (2024)
Continuous Reasoning for Vision-Language-Action
by: Wu, Yueh-Hua, et al.
Published: (2026)
by: Wu, Yueh-Hua, et al.
Published: (2026)
Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
by: Yuan, Xiu, et al.
Published: (2024)
by: Yuan, Xiu, et al.
Published: (2024)
Equivariant Action Sampling for Reinforcement Learning and Planning
by: Zhao, Linfeng, et al.
Published: (2024)
by: Zhao, Linfeng, et al.
Published: (2024)
Similar Items
-
When Maximum Entropy Misleads Policy Optimization
by: Zhang, Ruipeng, et al.
Published: (2025) -
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024) -
Improving Value Estimation Critically Enhances Vanilla Policy Gradient
by: Wang, Tao, et al.
Published: (2025) -
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
by: Zhang, Chen Bo Calvin, et al.
Published: (2024) -
Fractal Landscapes in Policy Optimization
by: Wang, Tao, et al.
Published: (2023)