Steering Your Diffusion Policy with Latent Space Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wagenmaker, Andrew, Nakamoto, Mitsuhiko, Zhang, Yunchu, Park, Seohong, Yagoub, Waleed, Nagabandi, Anusha, Gupta, Abhishek, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation
by: Su, Entong, et al.
Published: (2026)
by: Su, Entong, et al.
Published: (2026)
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
by: Hong, Matthew M., et al.
Published: (2026)
by: Hong, Matthew M., et al.
Published: (2026)
Foundation Policies with Hilbert Representations
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
Behavioral Exploration: Learning to Explore via In-Context Adaptation
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
by: Park, Seohong, et al.
Published: (2023)
by: Park, Seohong, et al.
Published: (2023)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
by: Park, Seohong, et al.
Published: (2023)
by: Park, Seohong, et al.
Published: (2023)
Decoupled Q-Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
Unsupervised-to-Online Reinforcement Learning
by: Kim, Junsu, et al.
Published: (2024)
by: Kim, Junsu, et al.
Published: (2024)
Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL
by: Wagenmaker, Andrew, et al.
Published: (2024)
by: Wagenmaker, Andrew, et al.
Published: (2024)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
Diffusion Guidance Is a Controllable Policy Improvement Operator
by: Frans, Kevin, et al.
Published: (2025)
by: Frans, Kevin, et al.
Published: (2025)
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
by: Patil, Sarvesh, et al.
Published: (2026)
by: Patil, Sarvesh, et al.
Published: (2026)
ASID: Active Exploration for System Identification in Robotic Manipulation
by: Memmel, Marius, et al.
Published: (2024)
by: Memmel, Marius, et al.
Published: (2024)
Latent-Conditioned Policy Gradient for Multi-Objective Deep Reinforcement Learning
by: Kanazawa, Takuya, et al.
Published: (2023)
by: Kanazawa, Takuya, et al.
Published: (2023)
RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
by: Xu, Charles, et al.
Published: (2024)
by: Xu, Charles, et al.
Published: (2024)
Flow Q-Learning
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
by: Frans, Kevin, et al.
Published: (2024)
by: Frans, Kevin, et al.
Published: (2024)
Symmetry-Aware Steering of Equivariant Diffusion Policies: Benefits and Limits
by: Park, Minwoo, et al.
Published: (2025)
by: Park, Minwoo, et al.
Published: (2025)
ATK: Automatic Task-driven Keypoint Selection for Robust Policy Learning
by: Zhang, Yunchu, et al.
Published: (2025)
by: Zhang, Yunchu, et al.
Published: (2025)
Latent Policy Steering through One-Step Flow Policies
by: Im, Hokyun, et al.
Published: (2026)
by: Im, Hokyun, et al.
Published: (2026)
Reinforcement Learning with Action Chunking
by: Li, Qiyang, et al.
Published: (2025)
by: Li, Qiyang, et al.
Published: (2025)
Multistep Quasimetric Learning for Scalable Goal-conditioned Reinforcement Learning
by: Zheng, Bill Chunyuan, et al.
Published: (2025)
by: Zheng, Bill Chunyuan, et al.
Published: (2025)
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
by: Lee, Tony, et al.
Published: (2026)
by: Lee, Tony, et al.
Published: (2026)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
CCIL: Continuity-based Data Augmentation for Corrective Imitation Learning
by: Ke, Liyiming, et al.
Published: (2023)
by: Ke, Liyiming, et al.
Published: (2023)
Dual Goal Representations
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024)
by: Doshi, Ria, et al.
Published: (2024)
From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
by: Wu, Yilin, et al.
Published: (2025)
by: Wu, Yilin, et al.
Published: (2025)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
by: Wang, Yiqi, et al.
Published: (2025)
by: Wang, Yiqi, et al.
Published: (2025)
Transitive RL: Value Learning via Divide and Conquer
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Is Value Learning Really the Main Bottleneck in Offline RL?
by: Park, Seohong, et al.
Published: (2024)
by: Park, Seohong, et al.
Published: (2024)
PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies
by: Jain, Arhan, et al.
Published: (2025)
by: Jain, Arhan, et al.
Published: (2025)
Real-Time Execution of Action Chunking Flow Policies
by: Black, Kevin, et al.
Published: (2025)
by: Black, Kevin, et al.
Published: (2025)
Scalable Offline Model-Based RL with Action Chunks
by: Park, Kwanyoung, et al.
Published: (2025)
by: Park, Kwanyoung, et al.
Published: (2025)
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
by: Yuan, Jessie, et al.
Published: (2026)
by: Yuan, Jessie, et al.
Published: (2026)
Policy Adaptation via Language Optimization: Decomposing Tasks for Few-Shot Imitation
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
Grammarization-Based Grasping with Deep Multi-Autoencoder Latent Space Exploration by Reinforcement Learning Agent
by: Askianakis, Leonidas
Published: (2024)
by: Askianakis, Leonidas
Published: (2024)
Similar Items
-
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
by: Nakamoto, Mitsuhiko, et al.
Published: (2024) -
RFS: Reinforcement Learning with Residual Flow Steering for Dexterous Manipulation
by: Su, Entong, et al.
Published: (2026) -
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
by: Hong, Matthew M., et al.
Published: (2026) -
Foundation Policies with Hilbert Representations
by: Park, Seohong, et al.
Published: (2024) -
Behavioral Exploration: Learning to Explore via In-Context Adaptation
by: Wagenmaker, Andrew, et al.
Published: (2025)