Residual Off-Policy RL for Finetuning Behavior Cloning Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Ankile, Lars, Jiang, Zhenyu, Duan, Rocky, Shi, Guanya, Abbeel, Pieter, Nagabandi, Anusha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
by: Hong, Matthew M., et al.
Published: (2026)
by: Hong, Matthew M., et al.
Published: (2026)
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
End-to-end RL Improves Dexterous Grasping Policies
by: Singh, Ritvik, et al.
Published: (2025)
by: Singh, Ritvik, et al.
Published: (2025)
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Diffusion Policy Policy Optimization
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
by: Yang, Lujie, et al.
Published: (2025)
by: Yang, Lujie, et al.
Published: (2025)
TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
by: Ze, Yanjie, et al.
Published: (2025)
by: Ze, Yanjie, et al.
Published: (2025)
Dataset Poisoning Attacks on Behavioral Cloning Policies
by: Kalra, Akansha, et al.
Published: (2025)
by: Kalra, Akansha, et al.
Published: (2025)
Safe Deep Policy Adaptation
by: Xiao, Wenli, et al.
Published: (2023)
by: Xiao, Wenli, et al.
Published: (2023)
TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint
by: Lin, Haotian, et al.
Published: (2025)
by: Lin, Haotian, et al.
Published: (2025)
Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching
by: Wu, Zhen, et al.
Published: (2026)
by: Wu, Zhen, et al.
Published: (2026)
Body Transformer: Leveraging Robot Embodiment for Policy Learning
by: Sferrazza, Carmelo, et al.
Published: (2024)
by: Sferrazza, Carmelo, et al.
Published: (2024)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies
by: Kalra, Akansha, et al.
Published: (2025)
by: Kalra, Akansha, et al.
Published: (2025)
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
by: Ye, Weirui, et al.
Published: (2025)
by: Ye, Weirui, et al.
Published: (2025)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
by: Sukhija, Bhavya, et al.
Published: (2024)
by: Sukhija, Bhavya, et al.
Published: (2024)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
DiffClone: Enhanced Behaviour Cloning in Robotics with Diffusion-Driven Policy Learning
by: Mani, Sabariswaran, et al.
Published: (2024)
by: Mani, Sabariswaran, et al.
Published: (2024)
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
by: Patil, Sarvesh, et al.
Published: (2026)
by: Patil, Sarvesh, et al.
Published: (2026)
Flow Policy Gradients for Robot Control
by: Yi, Brent, et al.
Published: (2026)
by: Yi, Brent, et al.
Published: (2026)
RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation
by: Su, Entong, et al.
Published: (2026)
by: Su, Entong, et al.
Published: (2026)
Offline Imitation Learning Through Graph Search and Retrieval
by: Yin, Zhao-Heng, et al.
Published: (2024)
by: Yin, Zhao-Heng, et al.
Published: (2024)
TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement Learning
by: Li, Ge, et al.
Published: (2024)
by: Li, Ge, et al.
Published: (2024)
How Generalizable Is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation
by: Vincent, Joseph A., et al.
Published: (2024)
by: Vincent, Joseph A., et al.
Published: (2024)
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
Robustifying a Policy in Multi-Agent RL with Diverse Cooperative Behaviors and Adversarial Style Sampling for Assistive Tasks
by: Osa, Takayuki, et al.
Published: (2024)
by: Osa, Takayuki, et al.
Published: (2024)
Refined Policy Distillation: From VLA Generalists to RL Experts
by: Jülg, Tobias, et al.
Published: (2025)
by: Jülg, Tobias, et al.
Published: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
by: Goli, Hossein, et al.
Published: (2025)
by: Goli, Hossein, et al.
Published: (2025)
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
by: Kim, Sungha, et al.
Published: (2026)
by: Kim, Sungha, et al.
Published: (2026)
Off Policy Lyapunov Stability in Reinforcement Learning
by: Gill, Sarvan, et al.
Published: (2025)
by: Gill, Sarvan, et al.
Published: (2025)
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning
by: Sun, Zhanyi, et al.
Published: (2026)
by: Sun, Zhanyi, et al.
Published: (2026)
Push Smarter, Not Harder: Hierarchical RL-Diffusion Policy for Efficient Nonprehensile Manipulation
by: Caro, Steven, et al.
Published: (2025)
by: Caro, Steven, et al.
Published: (2025)
Formulating Reinforcement Learning for Human-Robot Collaboration through Off-Policy Evaluation
by: Singh, Saurav, et al.
Published: (2026)
by: Singh, Saurav, et al.
Published: (2026)
Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics
by: Zhong, Yichao, et al.
Published: (2025)
by: Zhong, Yichao, et al.
Published: (2025)
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
by: Choi, Wonhyeok, et al.
Published: (2026)
by: Choi, Wonhyeok, et al.
Published: (2026)
Similar Items
-
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
by: Hong, Matthew M., et al.
Published: (2026) -
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
by: Zhao, Siheng, et al.
Published: (2025) -
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025) -
Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
by: Seo, Younggyo, et al.
Published: (2025) -
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)