REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Zhaoyuan, Chen, Yipu, Chai, Zimeng, Cueva, Alfred, Nguyen, Thong, Wu, Yifan, Xue, Huishu, Kim, Minji, Legene, Isaac, Liu, Fukang, Kim, Matthew, Barula, Ayan, Chen, Yongxin, Zhao, Ye
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918395034730496
author Gu, Zhaoyuan
Chen, Yipu
Chai, Zimeng
Cueva, Alfred
Nguyen, Thong
Wu, Yifan
Xue, Huishu
Kim, Minji
Legene, Isaac
Liu, Fukang
Kim, Matthew
Barula, Ayan
Chen, Yongxin
Zhao, Ye
author_facet Gu, Zhaoyuan
Chen, Yipu
Chai, Zimeng
Cueva, Alfred
Nguyen, Thong
Wu, Yifan
Xue, Huishu
Kim, Minji
Legene, Isaac
Liu, Fukang
Kim, Matthew
Barula, Ayan
Chen, Yongxin
Zhao, Ye
contents Humanoid loco-manipulation requires coordinated high-level motion plans with stable, low-level whole-body execution under complex robot-environment dynamics and long-horizon tasks. While diffusion policies (DPs) show promise for learning from demonstrations, deploying them on humanoids poses critical challenges: the motion planner trained offline is decoupled from the low-level controller, leading to poor command tracking, compounding distribution shift, and task failures. The common approach of scaling demonstration data is prohibitively expensive for high-dimensional humanoid systems. To address this challenge, we present REFINE-DP (REinforcement learning FINE-tuning of Diffusion Policy), a hierarchical framework that jointly optimizes a DP high-level planner and an RL-based low-level loco-manipulation controller. The DP is fine-tuned via a PPO-based diffusion policy gradient to improve task success rate, while the controller is simultaneously updated to accurately track the planner's evolving command distribution, reducing the distributional mismatch that degrades motion quality. We validate REFINE-DP on a humanoid robot performing loco-manipulation tasks, including door traversal and long-horizon object transport. REFINE-DP achieves an over $90\%$ success rate in simulation, even in out-of-distribution cases not seen in the pre-trained data, and enables smooth autonomous task execution in real-world dynamic environments. Our proposed method substantially outperforms pre-trained DP baselines and demonstrates that RL fine-tuning is key to reliable humanoid loco-manipulation. https://refine-dp.github.io/REFINE-DP/
format Preprint
id arxiv_https___arxiv_org_abs_2603_13707
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
Gu, Zhaoyuan
Chen, Yipu
Chai, Zimeng
Cueva, Alfred
Nguyen, Thong
Wu, Yifan
Xue, Huishu
Kim, Minji
Legene, Isaac
Liu, Fukang
Kim, Matthew
Barula, Ayan
Chen, Yongxin
Zhao, Ye
Robotics
Artificial Intelligence
Machine Learning
Humanoid loco-manipulation requires coordinated high-level motion plans with stable, low-level whole-body execution under complex robot-environment dynamics and long-horizon tasks. While diffusion policies (DPs) show promise for learning from demonstrations, deploying them on humanoids poses critical challenges: the motion planner trained offline is decoupled from the low-level controller, leading to poor command tracking, compounding distribution shift, and task failures. The common approach of scaling demonstration data is prohibitively expensive for high-dimensional humanoid systems. To address this challenge, we present REFINE-DP (REinforcement learning FINE-tuning of Diffusion Policy), a hierarchical framework that jointly optimizes a DP high-level planner and an RL-based low-level loco-manipulation controller. The DP is fine-tuned via a PPO-based diffusion policy gradient to improve task success rate, while the controller is simultaneously updated to accurately track the planner's evolving command distribution, reducing the distributional mismatch that degrades motion quality. We validate REFINE-DP on a humanoid robot performing loco-manipulation tasks, including door traversal and long-horizon object transport. REFINE-DP achieves an over $90\%$ success rate in simulation, even in out-of-distribution cases not seen in the pre-trained data, and enables smooth autonomous task execution in real-world dynamic environments. Our proposed method substantially outperforms pre-trained DP baselines and demonstrates that RL fine-tuning is key to reliable humanoid loco-manipulation. https://refine-dp.github.io/REFINE-DP/
title REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.13707