Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Xiu, Mu, Tongzhou, Tao, Stone, Fang, Yunhao, Zhang, Mengke, Su, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
von: Escoriza, Adrià López, et al.
Veröffentlicht: (2025)
von: Escoriza, Adrià López, et al.
Veröffentlicht: (2025)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
von: Wang, Yiqi, et al.
Veröffentlicht: (2025)
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024)
Sensory-Motor Control with Large Language Models via Iterative Policy Refinement
von: Carvalho, Jônata Tyska, et al.
Veröffentlicht: (2025)
von: Carvalho, Jônata Tyska, et al.
Veröffentlicht: (2025)
Harnessing Bounded-Support Evolution Strategies for Policy Refinement
von: Hirschowitz, Ethan, et al.
Veröffentlicht: (2025)
von: Hirschowitz, Ethan, et al.
Veröffentlicht: (2025)
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
von: Xue, Han, et al.
Veröffentlicht: (2025)
von: Xue, Han, et al.
Veröffentlicht: (2025)
Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning
von: Tao, Stone, et al.
Veröffentlicht: (2024)
von: Tao, Stone, et al.
Veröffentlicht: (2024)
Evolutionary Policy Optimization
von: Wang, Jianren, et al.
Veröffentlicht: (2025)
von: Wang, Jianren, et al.
Veröffentlicht: (2025)
Refining Compositional Diffusion for Reliable Long-Horizon Planning
von: Lee, Kyowoon, et al.
Veröffentlicht: (2026)
von: Lee, Kyowoon, et al.
Veröffentlicht: (2026)
Imagination Policy: Using Generative Point Cloud Models for Learning Manipulation Policies
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction
von: Jiang, Yunfan, et al.
Veröffentlicht: (2024)
von: Jiang, Yunfan, et al.
Veröffentlicht: (2024)
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
von: Zhang, Chen Bo Calvin, et al.
Veröffentlicht: (2024)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
von: Shitanda, Naoki, et al.
Veröffentlicht: (2026)
von: Shitanda, Naoki, et al.
Veröffentlicht: (2026)
Towards Embodiment Scaling Laws in Robot Locomotion
von: Ai, Bo, et al.
Veröffentlicht: (2025)
von: Ai, Bo, et al.
Veröffentlicht: (2025)
LearningFlow: Automated Policy Learning Workflow for Urban Driving with Large Language Models
von: Peng, Zengqi, et al.
Veröffentlicht: (2025)
von: Peng, Zengqi, et al.
Veröffentlicht: (2025)
Don't Start from Scratch: Behavioral Refinement via Interpolant-based Policy Diffusion
von: Chen, Kaiqi, et al.
Veröffentlicht: (2024)
von: Chen, Kaiqi, et al.
Veröffentlicht: (2024)
Mollification Effects of Policy Gradient Methods
von: Wang, Tao, et al.
Veröffentlicht: (2024)
von: Wang, Tao, et al.
Veröffentlicht: (2024)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
von: Cohen, Taco, et al.
Veröffentlicht: (2025)
Exploiting Hybrid Policy in Reinforcement Learning for Interpretable Temporal Logic Manipulation
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
Robot Policy Transfer with Online Demonstrations: An Active Reinforcement Learning Approach
von: Hou, Muhan, et al.
Veröffentlicht: (2025)
von: Hou, Muhan, et al.
Veröffentlicht: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
von: K, Swaminathan S, et al.
Veröffentlicht: (2026)
von: K, Swaminathan S, et al.
Veröffentlicht: (2026)
PWM: Policy Learning with Multi-Task World Models
von: Georgiev, Ignat, et al.
Veröffentlicht: (2024)
von: Georgiev, Ignat, et al.
Veröffentlicht: (2024)
Responsive Noise-Relaying Diffusion Policy: Responsive and Efficient Visuomotor Control
von: Chen, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Chen, Zhuoqun, et al.
Veröffentlicht: (2025)
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
von: Choi, Wonhyeok, et al.
Veröffentlicht: (2026)
von: Choi, Wonhyeok, et al.
Veröffentlicht: (2026)
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
von: Xiao, Wei, et al.
Veröffentlicht: (2025)
RLZero: Direct Policy Inference from Language Without In-Domain Supervision
von: Sikchi, Harshit, et al.
Veröffentlicht: (2024)
von: Sikchi, Harshit, et al.
Veröffentlicht: (2024)
ManiSkill-HAB: A Benchmark for Low-Level Manipulation in Home Rearrangement Tasks
von: Shukla, Arth, et al.
Veröffentlicht: (2024)
von: Shukla, Arth, et al.
Veröffentlicht: (2024)
RoboPocket: Improve Robot Policies Instantly with Your Phone
von: Fang, Junjie, et al.
Veröffentlicht: (2026)
von: Fang, Junjie, et al.
Veröffentlicht: (2026)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
von: Jin, Yang, et al.
Veröffentlicht: (2025)
von: Jin, Yang, et al.
Veröffentlicht: (2025)
Fisher Decorator: Refining Flow Policy via a Local Transport Map
von: Cheng, Xiaoyuan, et al.
Veröffentlicht: (2026)
von: Cheng, Xiaoyuan, et al.
Veröffentlicht: (2026)
Policy-Guided Diffusion
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
von: Jackson, Matthew Thomas, et al.
Veröffentlicht: (2024)
Absolute Policy Optimization
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
von: Zhao, Weiye, et al.
Veröffentlicht: (2023)
Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
von: Li, Chenhao, et al.
Veröffentlicht: (2025)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
von: Liu, Tenglong, et al.
Veröffentlicht: (2024)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
von: Lu, Dekun, et al.
Veröffentlicht: (2025)
von: Lu, Dekun, et al.
Veröffentlicht: (2025)
Towards Interpretable Foundation Models of Robot Behavior: A Task Specific Policy Generation Approach
von: Sheidlower, Isaac, et al.
Veröffentlicht: (2024)
von: Sheidlower, Isaac, et al.
Veröffentlicht: (2024)
Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
von: Cai, Shizhe, et al.
Veröffentlicht: (2025)
von: Cai, Shizhe, et al.
Veröffentlicht: (2025)
Foundation Policies with Hilbert Representations
von: Park, Seohong, et al.
Veröffentlicht: (2024)
von: Park, Seohong, et al.
Veröffentlicht: (2024)
Safe Deep Policy Adaptation
von: Xiao, Wenli, et al.
Veröffentlicht: (2023)
von: Xiao, Wenli, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DrS: Learning Reusable Dense Rewards for Multi-Stage Tasks
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024) -
Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning
von: Escoriza, Adrià López, et al.
Veröffentlicht: (2025) -
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
von: Wang, Yiqi, et al.
Veröffentlicht: (2025) -
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
von: Mu, Tongzhou, et al.
Veröffentlicht: (2024) -
Sensory-Motor Control with Large Language Models via Iterative Policy Refinement
von: Carvalho, Jônata Tyska, et al.
Veröffentlicht: (2025)