Efficient Robotic Policy Learning via Latent Space Backward Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Dongxiu, Niu, Haoyi, Wang, Zhihao, Zheng, Jinliang, Zheng, Yinan, Ou, Zhonghong, Hu, Jianming, Li, Jianxiong, Zhan, Xianyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916761452937216
author Liu, Dongxiu
Niu, Haoyi
Wang, Zhihao
Zheng, Jinliang
Zheng, Yinan
Ou, Zhonghong
Hu, Jianming
Li, Jianxiong
Zhan, Xianyuan
author_facet Liu, Dongxiu
Niu, Haoyi
Wang, Zhihao
Zheng, Jinliang
Zheng, Yinan
Ou, Zhonghong
Hu, Jianming
Li, Jianxiong
Zhan, Xianyuan
contents Current robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy learning: substantial computational costs that hinder real-time deployment, and accumulated inaccuracies that can mislead action extraction. Planning with coarse-grained subgoals partially alleviates efficiency issues. However, their forward planning schemes can still result in off-task predictions due to accumulation errors, leading to misalignment with long-term goals. This raises a critical question: Can robotic planning be both efficient and accurate enough for real-time control in long-horizon, multi-stage tasks? To address this, we propose a Latent Space Backward Planning scheme (LBP), which begins by grounding the task into final latent goals, followed by recursively predicting intermediate subgoals closer to the current state. The grounded final goal enables backward subgoal planning to always remain aware of task completion, facilitating on-task prediction along the entire planning horizon. The subgoal-conditioned policy incorporates a learnable token to summarize the subgoal sequences and determines how each subgoal guides action extraction. Through extensive simulation and real-robot long-horizon experiments, we show that LBP outperforms existing fine-grained and forward planning methods, achieving SOTA performance. Project Page: https://lbp-authors.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2505_06861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Robotic Policy Learning via Latent Space Backward Planning
Liu, Dongxiu
Niu, Haoyi
Wang, Zhihao
Zheng, Jinliang
Zheng, Yinan
Ou, Zhonghong
Hu, Jianming
Li, Jianxiong
Zhan, Xianyuan
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Current robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy learning: substantial computational costs that hinder real-time deployment, and accumulated inaccuracies that can mislead action extraction. Planning with coarse-grained subgoals partially alleviates efficiency issues. However, their forward planning schemes can still result in off-task predictions due to accumulation errors, leading to misalignment with long-term goals. This raises a critical question: Can robotic planning be both efficient and accurate enough for real-time control in long-horizon, multi-stage tasks? To address this, we propose a Latent Space Backward Planning scheme (LBP), which begins by grounding the task into final latent goals, followed by recursively predicting intermediate subgoals closer to the current state. The grounded final goal enables backward subgoal planning to always remain aware of task completion, facilitating on-task prediction along the entire planning horizon. The subgoal-conditioned policy incorporates a learnable token to summarize the subgoal sequences and determines how each subgoal guides action extraction. Through extensive simulation and real-robot long-horizon experiments, we show that LBP outperforms existing fine-grained and forward planning methods, achieving SOTA performance. Project Page: https://lbp-authors.github.io
title Efficient Robotic Policy Learning via Latent Space Backward Planning
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.06861