TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Matthew M., Zhang, Jesse, Nagabandi, Anusha, Gupta, Abhishek |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025)
by: Ankile, Lars, et al.
Published: (2025)
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
by: Wagenmaker, Andrew, et al.
Published: (2025)
by: Wagenmaker, Andrew, et al.
Published: (2025)
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
by: Zhu, Chuning, et al.
Published: (2025)
by: Zhu, Chuning, et al.
Published: (2025)
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)
by: Zhang, Jesse, et al.
Published: (2025)
Efficient Evaluation of Multi-Task Robot Policies With Active Experiment Selection
by: Anwar, Abrar, et al.
Published: (2025)
by: Anwar, Abrar, et al.
Published: (2025)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Policy-Guided Diffusion
by: Jackson, Matthew Thomas, et al.
Published: (2024)
by: Jackson, Matthew Thomas, et al.
Published: (2024)
Simulation Distillation: Pretraining World Models in Simulation for Rapid Real-World Adaptation
by: Levy, Jacob, et al.
Published: (2026)
by: Levy, Jacob, et al.
Published: (2026)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
by: Di Palo, Norman, et al.
Published: (2024)
by: Di Palo, Norman, et al.
Published: (2024)
TAIL: Task-specific Adapters for Imitation Learning with Large Pretrained Models
by: Liu, Zuxin, et al.
Published: (2023)
by: Liu, Zuxin, et al.
Published: (2023)
EXTRACT: Efficient Policy Learning by Extracting Transferable Robot Skills from Offline Data
by: Zhang, Jesse, et al.
Published: (2024)
by: Zhang, Jesse, et al.
Published: (2024)
Latent Policy Steering with Embodiment-Agnostic Pretrained World Models
by: Wang, Yiqi, et al.
Published: (2025)
by: Wang, Yiqi, et al.
Published: (2025)
Safe Exploration via Policy Priors
by: Wendl, Manuel, et al.
Published: (2026)
by: Wendl, Manuel, et al.
Published: (2026)
Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
by: Yokozawa, Riko, et al.
Published: (2025)
by: Yokozawa, Riko, et al.
Published: (2025)
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
by: Pappalardo, Octavio
Published: (2026)
by: Pappalardo, Octavio
Published: (2026)
SPRINT: Scalable Policy Pre-Training via Language Instruction Relabeling
by: Zhang, Jesse, et al.
Published: (2023)
by: Zhang, Jesse, et al.
Published: (2023)
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies
by: Patil, Sarvesh, et al.
Published: (2026)
by: Patil, Sarvesh, et al.
Published: (2026)
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
VAMOS: A Hierarchical Vision-Language-Action Model for Capability-Modulated and Steerable Navigation
by: Castro, Mateo Guaman, et al.
Published: (2025)
by: Castro, Mateo Guaman, et al.
Published: (2025)
Evolutionary Policy Optimization
by: Wang, Jianren, et al.
Published: (2025)
by: Wang, Jianren, et al.
Published: (2025)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
by: Xue, Han, et al.
Published: (2025)
by: Xue, Han, et al.
Published: (2025)
Fine-tuning Diffusion Policies with Backpropagation Through Diffusion Timesteps
by: Yang, Ningyuan, et al.
Published: (2025)
by: Yang, Ningyuan, et al.
Published: (2025)
Variational Distillation of Diffusion Policies into Mixture of Experts
by: Zhou, Hongyi, et al.
Published: (2024)
by: Zhou, Hongyi, et al.
Published: (2024)
Adaptive Diffusion Policy Optimization for Robotic Manipulation
by: Jiang, Huiyun, et al.
Published: (2025)
by: Jiang, Huiyun, et al.
Published: (2025)
ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
by: Zhang, Yunuo, et al.
Published: (2025)
by: Zhang, Yunuo, et al.
Published: (2025)
TopSpark: A Timestep Optimization Methodology for Energy-Efficient Spiking Neural Networks on Autonomous Mobile Agents
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2023)
by: Putra, Rachmad Vidya Wicaksana, et al.
Published: (2023)
Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions
by: Hao, Alexis Yihong, et al.
Published: (2025)
by: Hao, Alexis Yihong, et al.
Published: (2025)
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss
by: Zheng, Ruijie, et al.
Published: (2024)
by: Zheng, Ruijie, et al.
Published: (2024)
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
by: Gu, Zhaoyuan, et al.
Published: (2026)
by: Gu, Zhaoyuan, et al.
Published: (2026)
Semantic World Models
by: Berg, Jacob, et al.
Published: (2025)
by: Berg, Jacob, et al.
Published: (2025)
Robust Policy Learning via Offline Skill Diffusion
by: Kim, Woo Kyung, et al.
Published: (2024)
by: Kim, Woo Kyung, et al.
Published: (2024)
WARPD: World model Assisted Reactive Policy Diffusion
by: Hegde, Shashank, et al.
Published: (2024)
by: Hegde, Shashank, et al.
Published: (2024)
Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design
by: Schlaginhaufen, Andreas, et al.
Published: (2025)
by: Schlaginhaufen, Andreas, et al.
Published: (2025)
In-Context Policy Adaptation via Cross-Domain Skill Diffusion
by: Yoo, Minjong, et al.
Published: (2025)
by: Yoo, Minjong, et al.
Published: (2025)
Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter Efficient
by: Wang, Wenlong, et al.
Published: (2024)
by: Wang, Wenlong, et al.
Published: (2024)
Similar Items
-
Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
by: Wagenmaker, Andrew, et al.
Published: (2025) -
Residual Off-Policy RL for Finetuning Behavior Cloning Policies
by: Ankile, Lars, et al.
Published: (2025) -
Steering Your Diffusion Policy with Latent Space Reinforcement Learning
by: Wagenmaker, Andrew, et al.
Published: (2025) -
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
by: Zhu, Chuning, et al.
Published: (2025) -
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
by: Zhang, Jesse, et al.
Published: (2025)