Long-Horizon Visual Imitation Learning via Plan and Code Reflection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Quan, Shi, Chenrui, Chen, Qi, Wu, Yuwei, Gao, Zhi, Zhang, Xintong, Gao, Rui, Wu, Kun, Jia, Yunde
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908719464316928
author Chen, Quan
Shi, Chenrui
Chen, Qi
Wu, Yuwei
Gao, Zhi
Zhang, Xintong
Gao, Rui
Wu, Kun
Jia, Yunde
author_facet Chen, Quan
Shi, Chenrui
Chen, Qi
Wu, Yuwei
Gao, Zhi
Zhang, Xintong
Gao, Rui
Wu, Kun
Jia, Yunde
contents Learning from long-horizon demonstrations with complex action sequences presents significant challenges for visual imitation learning, particularly in understanding temporal relationships of actions and spatial relationships between objects. In this paper, we propose a new agent framework that incorporates two dedicated reflection modules to enhance both plan and code generation. The plan generation module produces an initial action sequence, which is then verified by the plan reflection module to ensure temporal coherence and spatial alignment with the demonstration video. The code generation module translates the plan into executable code, while the code reflection module verifies and refines the generated code to ensure correctness and consistency with the generated plan. These two reflection modules jointly enable the agent to detect and correct errors in both the plan generation and code generation, improving performance in tasks with intricate temporal and spatial dependencies. To support systematic evaluation, we introduce LongVILBench, a benchmark comprising 300 human demonstrations with action sequences of up to 18 steps. LongVILBench emphasizes temporal and spatial complexity across multiple task types. Experimental results demonstrate that existing methods perform poorly on this benchmark, whereas our new framework establishes a strong baseline for long-horizon visual imitation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long-Horizon Visual Imitation Learning via Plan and Code Reflection
Chen, Quan
Shi, Chenrui
Chen, Qi
Wu, Yuwei
Gao, Zhi
Zhang, Xintong
Gao, Rui
Wu, Kun
Jia, Yunde
Robotics
Artificial Intelligence
Machine Learning
I.2.9; I.2.10
Learning from long-horizon demonstrations with complex action sequences presents significant challenges for visual imitation learning, particularly in understanding temporal relationships of actions and spatial relationships between objects. In this paper, we propose a new agent framework that incorporates two dedicated reflection modules to enhance both plan and code generation. The plan generation module produces an initial action sequence, which is then verified by the plan reflection module to ensure temporal coherence and spatial alignment with the demonstration video. The code generation module translates the plan into executable code, while the code reflection module verifies and refines the generated code to ensure correctness and consistency with the generated plan. These two reflection modules jointly enable the agent to detect and correct errors in both the plan generation and code generation, improving performance in tasks with intricate temporal and spatial dependencies. To support systematic evaluation, we introduce LongVILBench, a benchmark comprising 300 human demonstrations with action sequences of up to 18 steps. LongVILBench emphasizes temporal and spatial complexity across multiple task types. Experimental results demonstrate that existing methods perform poorly on this benchmark, whereas our new framework establishes a strong baseline for long-horizon visual imitation learning.
title Long-Horizon Visual Imitation Learning via Plan and Code Reflection
topic Robotics
Artificial Intelligence
Machine Learning
I.2.9; I.2.10
url https://arxiv.org/abs/2509.05368