Learning Generalizable Visuomotor Policy through Dynamics-Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Dohyeok, Lee, Jung Min, Kim, Munkyung, Ju, Seokhun, Koo, Jin Woo, Lee, Kyungjae, Kim, Dohyeong, Cho, TaeHyun, Lee, Jungwoo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918180312580096
author Lee, Dohyeok
Lee, Jung Min
Kim, Munkyung
Ju, Seokhun
Koo, Jin Woo
Lee, Kyungjae
Kim, Dohyeong
Cho, TaeHyun
Lee, Jungwoo
author_facet Lee, Dohyeok
Lee, Jung Min
Kim, Munkyung
Ju, Seokhun
Koo, Jin Woo
Lee, Kyungjae
Kim, Dohyeong
Cho, TaeHyun
Lee, Jungwoo
contents Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction models have shown promising results by learning rich spatiotemporal representations from large-scale datasets. However, these models learn action-agnostic dynamics that cannot distinguish between different control inputs, limiting their utility for precise manipulation tasks and requiring large pretraining datasets. We propose a Dynamics-Aligned Flow Matching Policy (DAP) that integrates dynamics prediction into policy learning. Our method introduces a novel architecture where policy and dynamics models provide mutual corrective feedback during action generation, enabling self-correction and improved generalization. Empirical validation demonstrates generalization performance superior to baseline methods on real-world robotic manipulation tasks, showing particular robustness in OOD scenarios including visual distractions and lighting variations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27114
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Generalizable Visuomotor Policy through Dynamics-Alignment
Lee, Dohyeok
Lee, Jung Min
Kim, Munkyung
Ju, Seokhun
Koo, Jin Woo
Lee, Kyungjae
Kim, Dohyeong
Cho, TaeHyun
Lee, Jungwoo
Robotics
Machine Learning
Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction models have shown promising results by learning rich spatiotemporal representations from large-scale datasets. However, these models learn action-agnostic dynamics that cannot distinguish between different control inputs, limiting their utility for precise manipulation tasks and requiring large pretraining datasets. We propose a Dynamics-Aligned Flow Matching Policy (DAP) that integrates dynamics prediction into policy learning. Our method introduces a novel architecture where policy and dynamics models provide mutual corrective feedback during action generation, enabling self-correction and improved generalization. Empirical validation demonstrates generalization performance superior to baseline methods on real-world robotic manipulation tasks, showing particular robustness in OOD scenarios including visual distractions and lighting variations.
title Learning Generalizable Visuomotor Policy through Dynamics-Alignment
topic Robotics
Machine Learning
url https://arxiv.org/abs/2510.27114