UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception
Fuente:
arXiv
Saved in:
| Main Author: | Wang, Ziming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset
by: Liu, Kehui, et al.
Published: (2025)
by: Liu, Kehui, et al.
Published: (2025)
UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
by: Rayyan, Omar, et al.
Published: (2025)
by: Rayyan, Omar, et al.
Published: (2025)
No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
by: Yu, Run, et al.
Published: (2025)
by: Yu, Run, et al.
Published: (2025)
EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation
by: Bode, Jonas, et al.
Published: (2025)
by: Bode, Jonas, et al.
Published: (2025)
BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
by: Li, Peiyan, et al.
Published: (2025)
by: Li, Peiyan, et al.
Published: (2025)
FP3: A 3D Foundation Policy for Robotic Manipulation
by: Yang, Rujia, et al.
Published: (2025)
by: Yang, Rujia, et al.
Published: (2025)
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
by: Li, Yuelei, et al.
Published: (2025)
by: Li, Yuelei, et al.
Published: (2025)
DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation
by: Xu, Mengda, et al.
Published: (2025)
by: Xu, Mengda, et al.
Published: (2025)
Dynamic Manipulation of Deformable Objects in 3D: Simulation, Benchmark and Learning Strategy
by: Lan, Guanzhou, et al.
Published: (2025)
by: Lan, Guanzhou, et al.
Published: (2025)
GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
by: Qian, Quanhao, et al.
Published: (2025)
by: Qian, Quanhao, et al.
Published: (2025)
A3D: Adaptive Affordance Assembly with Dual-Arm Manipulation
by: Liang, Jiaqi, et al.
Published: (2026)
by: Liang, Jiaqi, et al.
Published: (2026)
FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset
by: Zhaxizhuoma, et al.
Published: (2024)
by: Zhaxizhuoma, et al.
Published: (2024)
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation
by: Zhao, Ziyan, et al.
Published: (2025)
by: Zhao, Ziyan, et al.
Published: (2025)
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
by: Lu, Guanxing, et al.
Published: (2024)
by: Lu, Guanxing, et al.
Published: (2024)
Manipulating Elasto-Plastic Objects With 3D Occupancy and Learning-Based Predictive Control
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
by: Dong, Zibin, et al.
Published: (2025)
by: Dong, Zibin, et al.
Published: (2025)
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
by: Xu, Tianling, et al.
Published: (2025)
by: Xu, Tianling, et al.
Published: (2025)
DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation
by: Lyu, Jiangran, et al.
Published: (2025)
by: Lyu, Jiangran, et al.
Published: (2025)
Flow as the Cross-Domain Manipulation Interface
by: Xu, Mengda, et al.
Published: (2024)
by: Xu, Mengda, et al.
Published: (2024)
MRHER: Model-based Relay Hindsight Experience Replay for Sequential Object Manipulation Tasks with Sparse Rewards
by: Huang, Yuming, et al.
Published: (2023)
by: Huang, Yuming, et al.
Published: (2023)
Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement
by: Malik, Ashish, et al.
Published: (2026)
by: Malik, Ashish, et al.
Published: (2026)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
by: Huang, Binghao, et al.
Published: (2024)
by: Huang, Binghao, et al.
Published: (2024)
Indoor and Outdoor 3D Scene Graph Generation via Language-Enabled Spatial Ontologies
by: Strader, Jared, et al.
Published: (2023)
by: Strader, Jared, et al.
Published: (2023)
Active-Perceptive Motion Generation for Mobile Manipulation
by: Jauhri, Snehal, et al.
Published: (2023)
by: Jauhri, Snehal, et al.
Published: (2023)
DeformerNet: Learning Bimanual Manipulation of 3D Deformable Objects
by: Thach, Bao, et al.
Published: (2023)
by: Thach, Bao, et al.
Published: (2023)
TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks
by: Cheng, Tailai, et al.
Published: (2026)
by: Cheng, Tailai, et al.
Published: (2026)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
by: Liu, Ruixun, et al.
Published: (2025)
by: Liu, Ruixun, et al.
Published: (2025)
3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting
by: Zheng, Wancai, et al.
Published: (2026)
by: Zheng, Wancai, et al.
Published: (2026)
RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
by: Liu, Songming, et al.
Published: (2026)
by: Liu, Songming, et al.
Published: (2026)
Unifying 2D and 3D Vision-Language Understanding
by: Jain, Ayush, et al.
Published: (2025)
by: Jain, Ayush, et al.
Published: (2025)
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
by: Eisner, Ben, et al.
Published: (2022)
by: Eisner, Ben, et al.
Published: (2022)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
3DSGrasp: 3D Shape-Completion for Robotic Grasp
by: Mohammadi, Seyed S., et al.
Published: (2023)
by: Mohammadi, Seyed S., et al.
Published: (2023)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning
by: Tang, Guoqin, et al.
Published: (2025)
by: Tang, Guoqin, et al.
Published: (2025)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
by: Cui, Wenbo, et al.
Published: (2025)
by: Cui, Wenbo, et al.
Published: (2025)
Point2Graph: An End-to-end Point Cloud-based 3D Open-Vocabulary Scene Graph for Robot Navigation
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
Similar Items
-
FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset
by: Liu, Kehui, et al.
Published: (2025) -
UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation
by: Li, Hao, et al.
Published: (2026) -
MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
by: Rayyan, Omar, et al.
Published: (2025) -
No Need for Real 3D: Fusing 2D Vision with Pseudo 3D Representations for Robotic Manipulation Learning
by: Yu, Run, et al.
Published: (2025) -
EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation
by: Bode, Jonas, et al.
Published: (2025)