DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Jusuk, Lee, Seungjae, Shin, Jonghun, Jung, Hoseong, Kim, Sungha, Cho, Daesol, Kim, H. Jin, Huang, Jia-Bin, Huang, Furong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910270509547520
author Lee, Jusuk
Lee, Seungjae
Shin, Jonghun
Jung, Hoseong
Kim, Sungha
Cho, Daesol
Kim, H. Jin
Huang, Jia-Bin
Huang, Furong
author_facet Lee, Jusuk
Lee, Seungjae
Shin, Jonghun
Jung, Hoseong
Kim, Sungha
Cho, Daesol
Kim, H. Jin
Huang, Jia-Bin
Huang, Furong
contents Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30350
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Lee, Jusuk
Lee, Seungjae
Shin, Jonghun
Jung, Hoseong
Kim, Sungha
Cho, Daesol
Kim, H. Jin
Huang, Jia-Bin
Huang, Furong
Robotics
Machine Learning
Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.
title DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
topic Robotics
Machine Learning
url https://arxiv.org/abs/2605.30350