DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910270509547520 |
|---|---|
| author | Lee, Jusuk Lee, Seungjae Shin, Jonghun Jung, Hoseong Kim, Sungha Cho, Daesol Kim, H. Jin Huang, Jia-Bin Huang, Furong |
| author_facet | Lee, Jusuk Lee, Seungjae Shin, Jonghun Jung, Hoseong Kim, Sungha Cho, Daesol Kim, H. Jin Huang, Jia-Bin Huang, Furong |
| contents | Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_30350 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation Lee, Jusuk Lee, Seungjae Shin, Jonghun Jung, Hoseong Kim, Sungha Cho, Daesol Kim, H. Jin Huang, Jia-Bin Huang, Furong Robotics Machine Learning Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action. |
| title | DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation |
| topic | Robotics Machine Learning |
| url | https://arxiv.org/abs/2605.30350 |