BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Rutav, Yu, Albert, Zhu, Yifeng, Zhu, Yuke, Martín-Martín, Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
by: Liu, Huihan, et al.
Published: (2025)
by: Liu, Huihan, et al.
Published: (2025)
LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery
by: Wan, Weikang, et al.
Published: (2023)
by: Wan, Weikang, et al.
Published: (2023)
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Efficient Sensorimotor Learning for Open-world Robot Manipulation
by: Zhu, Yifeng
Published: (2025)
by: Zhu, Yifeng
Published: (2025)
Model-Based Runtime Monitoring with Interactive Imitation Learning
by: Liu, Huihan, et al.
Published: (2023)
by: Liu, Huihan, et al.
Published: (2023)
Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
by: Jiang, Zhenyu, et al.
Published: (2024)
by: Jiang, Zhenyu, et al.
Published: (2024)
MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning
by: Mohammadi, Mohammad, et al.
Published: (2025)
by: Mohammadi, Mohammad, et al.
Published: (2025)
MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
by: Shah, Rutav, et al.
Published: (2025)
by: Shah, Rutav, et al.
Published: (2025)
SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation
by: Bahety, Arpit, et al.
Published: (2025)
by: Bahety, Arpit, et al.
Published: (2025)
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
by: Zhu, Yuke, et al.
Published: (2020)
by: Zhu, Yuke, et al.
Published: (2020)
Acting and Planning with Hierarchical Operational Models on a Mobile Robot: A Study with RAE+UPOM
by: Lima, Oscar, et al.
Published: (2025)
by: Lima, Oscar, et al.
Published: (2025)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
by: Liu, Huihan, et al.
Published: (2026)
by: Liu, Huihan, et al.
Published: (2026)
RoboSSM: Scalable In-context Imitation Learning via State-Space Models
by: Yoo, Youngju, et al.
Published: (2025)
by: Yoo, Youngju, et al.
Published: (2025)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
by: Takagi, Yusuke, et al.
Published: (2026)
by: Takagi, Yusuke, et al.
Published: (2026)
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
by: Li, Jinhan, et al.
Published: (2024)
by: Li, Jinhan, et al.
Published: (2024)
Premover: Fast Vision-Language-Action Control by Acting Before Instructions Are Complete
by: Park, Joonha, et al.
Published: (2026)
by: Park, Joonha, et al.
Published: (2026)
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
by: Jin, Ruixing, et al.
Published: (2026)
by: Jin, Ruixing, et al.
Published: (2026)
CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model
by: Wang, Feiyang, et al.
Published: (2025)
by: Wang, Feiyang, et al.
Published: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
PRIME: Scaffolding Manipulation Tasks with Behavior Primitives for Data-Efficient Imitation Learning
by: Gao, Tian, et al.
Published: (2024)
by: Gao, Tian, et al.
Published: (2024)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
by: Su, Taiyi, et al.
Published: (2026)
by: Su, Taiyi, et al.
Published: (2026)
Zero-shot Object Navigation with Vision-Language Models Reasoning
by: Wen, Congcong, et al.
Published: (2024)
by: Wen, Congcong, et al.
Published: (2024)
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
by: Hong, Youngjin, et al.
Published: (2025)
by: Hong, Youngjin, et al.
Published: (2025)
KinScene: Model-Based Mobile Manipulation of Articulated Scenes
by: Hsu, Cheng-Chun, et al.
Published: (2024)
by: Hsu, Cheng-Chun, et al.
Published: (2024)
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
by: Gubernatorov, Konstantin, et al.
Published: (2025)
by: Gubernatorov, Konstantin, et al.
Published: (2025)
TeleMoMa: A Modular and Versatile Teleoperation System for Mobile Manipulation
by: Dass, Shivin, et al.
Published: (2024)
by: Dass, Shivin, et al.
Published: (2024)
Build on Priors: Vision--Language--Guided Neuro-Symbolic Imitation Learning for Data-Efficient Real-World Robot Manipulation
by: Lorang, Pierrick, et al.
Published: (2026)
by: Lorang, Pierrick, et al.
Published: (2026)
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
by: Lin, Kevin, et al.
Published: (2026)
by: Lin, Kevin, et al.
Published: (2026)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
by: Lin, Toru, et al.
Published: (2025)
by: Lin, Toru, et al.
Published: (2025)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025)
by: Kawaharazuka, Kento, et al.
Published: (2025)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation
by: Maddukuri, Abhiram, et al.
Published: (2025)
by: Maddukuri, Abhiram, et al.
Published: (2025)
Multi-Task Interactive Robot Fleet Learning with Visual World Models
by: Liu, Huihan, et al.
Published: (2024)
by: Liu, Huihan, et al.
Published: (2024)
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
by: Zhu, Chuning, et al.
Published: (2025)
by: Zhu, Chuning, et al.
Published: (2025)
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
by: Chen, Annie S., et al.
Published: (2024)
by: Chen, Annie S., et al.
Published: (2024)
LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
by: Zhu, Fangqi, et al.
Published: (2025)
by: Zhu, Fangqi, et al.
Published: (2025)
SKT: Integrating State-Aware Keypoint Trajectories with Vision-Language Models for Robotic Garment Manipulation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Similar Items
-
Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models
by: Liu, Huihan, et al.
Published: (2025) -
LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery
by: Wan, Weikang, et al.
Published: (2023) -
SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning
by: Zhang, Yu, et al.
Published: (2025) -
Efficient Sensorimotor Learning for Open-world Robot Manipulation
by: Zhu, Yifeng
Published: (2025) -
Model-Based Runtime Monitoring with Interactive Imitation Learning
by: Liu, Huihan, et al.
Published: (2023)