RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Chengkai, Wu, Kun, Liu, Jiaming, Che, Zhengping, Wu, Di, Liao, Fei, Li, Guangrun, He, Jingyang, Feng, Qiuxuan, Jin, Zhao, Gu, Chenyang, Liu, Zhuoyang, Han, Nuowei, Mi, Xiangju, Lv, Yaoxu, Fu, Yankai, Dai, Gaole, Gu, Langzhe, Li, Tao, Zhang, Yuheng, Zhang, Yixue, Wang, Xinhua, Fan, Shichao, Li, Meng, Zhao, Zhen, Liu, Ning, Xu, Zhiyuan, Ren, Pei, Ji, Junjie, Liu, Haonan, Cheng, Kuan, Zhang, Shanghang, Tang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
by: Li, Guangrun, et al.
Published: (2025)
by: Li, Guangrun, et al.
Published: (2025)
Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation
by: He, Jingyang, et al.
Published: (2026)
by: He, Jingyang, et al.
Published: (2026)
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
by: Zhang, Yixue, et al.
Published: (2026)
by: Zhang, Yixue, et al.
Published: (2026)
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
by: Liu, Zhuoyang, et al.
Published: (2026)
by: Liu, Zhuoyang, et al.
Published: (2026)
RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
by: Wang, Xinhua, et al.
Published: (2026)
by: Wang, Xinhua, et al.
Published: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets
by: Wu, Zhuangzhe, et al.
Published: (2026)
by: Wu, Zhuangzhe, et al.
Published: (2026)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
by: Fan, Shichao, et al.
Published: (2025)
by: Fan, Shichao, et al.
Published: (2025)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
by: Luo, Yulin, et al.
Published: (2026)
by: Luo, Yulin, et al.
Published: (2026)
RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
by: Wu, Shihan, et al.
Published: (2025)
by: Wu, Shihan, et al.
Published: (2025)
AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation
by: Chen, Sixiang, et al.
Published: (2025)
by: Chen, Sixiang, et al.
Published: (2025)
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
by: Xu, Qinwen, et al.
Published: (2026)
by: Xu, Qinwen, et al.
Published: (2026)
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
by: Fu, Yankai, et al.
Published: (2025)
by: Fu, Yankai, et al.
Published: (2025)
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
Multimodal Large Language Models for Bioimage Analysis
by: Zhang, Shanghang, et al.
Published: (2024)
by: Zhang, Shanghang, et al.
Published: (2024)
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
by: Gu, Chenyang, et al.
Published: (2025)
by: Gu, Chenyang, et al.
Published: (2025)
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
by: Feng, Qiuxuan, et al.
Published: (2026)
by: Feng, Qiuxuan, et al.
Published: (2026)
Orochi: Versatile Biomedical Image Processor
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Chinese women faculty in US universities: Transnational experiences and challenges
by: Xiangju Liu, et al.
Published: (2024)
by: Xiangju Liu, et al.
Published: (2024)
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
by: Zhang, Hantao, et al.
Published: (2025)
by: Zhang, Hantao, et al.
Published: (2025)
URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop
by: Chen, Tianxing, et al.
Published: (2025)
by: Chen, Tianxing, et al.
Published: (2025)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
by: Zhang, Zongzheng, et al.
Published: (2026)
by: Zhang, Zongzheng, et al.
Published: (2026)
Investigation on the Dynamic Behavior of Rock in Three Stages of Rock Fragmentation Under Vibro‐Impact Condition: Continuum and Discontinuum Model
by: Zhao Zhang, et al.
Published: (2025)
by: Zhao Zhang, et al.
Published: (2025)
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
by: Fan, Shichao, et al.
Published: (2025)
by: Fan, Shichao, et al.
Published: (2025)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
MIND: Benchmarking Memory Consistency and Action Control in World Models
by: Ye, Yixuan, et al.
Published: (2026)
by: Ye, Yixuan, et al.
Published: (2026)
RoboBrain 2.0 Technical Report
by: BAAI RoboBrain Team, et al.
Published: (2025)
by: BAAI RoboBrain Team, et al.
Published: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
by: Chen, Tianxing, et al.
Published: (2025)
by: Chen, Tianxing, et al.
Published: (2025)
Spatiotemporal correlation and joint coverage probability in ultra‐dense mobile networks
by: Zhuoyang Li, et al.
Published: (2024)
by: Zhuoyang Li, et al.
Published: (2024)
VTouch++: A Multimodal Dataset with Vision-Based Tactile Enhancement for Bimanual Manipulation
by: Hua, Qianxi, et al.
Published: (2026)
by: Hua, Qianxi, et al.
Published: (2026)
Similar Items
-
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
by: Wu, Kun, et al.
Published: (2024) -
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
by: Li, Guangrun, et al.
Published: (2025) -
Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation
by: He, Jingyang, et al.
Published: (2026) -
RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation
by: Zhang, Yixue, et al.
Published: (2026) -
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
by: Liu, Zhuoyang, et al.
Published: (2026)