Towards Generalizable Robotic Manipulation in Dynamic Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Heng, Li, Shangru, Wang, Shuhan, Xi, Xuanyang, Liang, Dingkang, Bai, Xiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911595101159424
author Fang, Heng
Li, Shangru
Wang, Shuhan
Xi, Xuanyang
Liang, Dingkang
Bai, Xiang
author_facet Fang, Heng
Li, Shangru
Wang, Shuhan
Xi, Xuanyang
Liang, Dingkang
Bai, Xiang
contents Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi-dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics-aware VLA architecture. By integrating scene-centric historical optical flow and specialized world queries to implicitly forecast object-centric future states, PUMA couples history-aware perception with short-horizon prediction. Results demonstrate that PUMA achieves state-of-the-art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H-EmbodVis/DOMINO.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15620
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Generalizable Robotic Manipulation in Dynamic Environments
Fang, Heng
Li, Shangru
Wang, Shuhan
Xi, Xuanyang
Liang, Dingkang
Bai, Xiang
Computer Vision and Pattern Recognition
Robotics
Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap primarily stems from a scarcity of dynamic manipulation datasets and the reliance of mainstream VLAs on single-frame observations, restricting their spatiotemporal reasoning capabilities. To address this, we introduce DOMINO, a large-scale dataset and benchmark for generalizable dynamic manipulation, featuring 35 tasks with hierarchical complexities, over 110K expert trajectories, and a multi-dimensional evaluation suite. Through comprehensive experiments, we systematically evaluate existing VLAs on dynamic tasks, explore effective training strategies for dynamic awareness, and validate the generalizability of dynamic data. Furthermore, we propose PUMA, a dynamics-aware VLA architecture. By integrating scene-centric historical optical flow and specialized world queries to implicitly forecast object-centric future states, PUMA couples history-aware perception with short-horizon prediction. Results demonstrate that PUMA achieves state-of-the-art performance, yielding a 6.3% absolute improvement in success rate over baselines. Moreover, we show that training on dynamic data fosters robust spatiotemporal representations that transfer to static tasks. All code and data are available at https://github.com/H-EmbodVis/DOMINO.
title Towards Generalizable Robotic Manipulation in Dynamic Environments
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2603.15620