Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Qi, Fu, Zipeng, Cheng, Xuxin, Wang, Xiaolong, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Playful DoggyBot: Learning Agile and Precise Quadrupedal Locomotion
by: Duan, Xin, et al.
Published: (2024)
by: Duan, Xin, et al.
Published: (2024)
FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
by: Liu, Weiheng, et al.
Published: (2025)
by: Liu, Weiheng, et al.
Published: (2025)
HumanPlus: Humanoid Shadowing and Imitation from Humans
by: Fu, Zipeng, et al.
Published: (2024)
by: Fu, Zipeng, et al.
Published: (2024)
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
by: Fu, Zipeng, et al.
Published: (2024)
by: Fu, Zipeng, et al.
Published: (2024)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
by: Cheng, An-Chieh, et al.
Published: (2024)
by: Cheng, An-Chieh, et al.
Published: (2024)
Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation
by: Jin, Qunchao, et al.
Published: (2026)
by: Jin, Qunchao, et al.
Published: (2026)
Visual Whole-Body Control for Legged Loco-Manipulation
by: Liu, Minghuan, et al.
Published: (2024)
by: Liu, Minghuan, et al.
Published: (2024)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation
by: Gong, Ruiyan, et al.
Published: (2026)
by: Gong, Ruiyan, et al.
Published: (2026)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
LOVON: Legged Open-Vocabulary Object Navigator
by: Peng, Daojie, et al.
Published: (2025)
by: Peng, Daojie, et al.
Published: (2025)
Towards Open-World Grasping with Large Vision-Language Models
by: Tziafas, Georgios, et al.
Published: (2024)
by: Tziafas, Georgios, et al.
Published: (2024)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
by: Liu, Fei, et al.
Published: (2025)
by: Liu, Fei, et al.
Published: (2025)
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
by: Wu, Jimmy, et al.
Published: (2024)
by: Wu, Jimmy, et al.
Published: (2024)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation
by: Nie, Dujun, et al.
Published: (2025)
by: Nie, Dujun, et al.
Published: (2025)
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs
by: Qiao, Yanyuan, et al.
Published: (2024)
by: Qiao, Yanyuan, et al.
Published: (2024)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
by: Zhao, Guoyang, et al.
Published: (2025)
by: Zhao, Guoyang, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
by: Xiao, Jianqiang, et al.
Published: (2025)
by: Xiao, Jianqiang, et al.
Published: (2025)
Vision-based Manipulation from Single Human Video with Open-World Object Graphs
by: Zhu, Yifeng, et al.
Published: (2024)
by: Zhu, Yifeng, et al.
Published: (2024)
Category-Level and Open-Set Object Pose Estimation for Robotics
by: Hönig, Peter, et al.
Published: (2025)
by: Hönig, Peter, et al.
Published: (2025)
OW-Rep: Open World Object Detection with Instance Representation Learning
by: Lee, Sunoh, et al.
Published: (2024)
by: Lee, Sunoh, et al.
Published: (2024)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
by: Zeng, Shuang, et al.
Published: (2025)
by: Zeng, Shuang, et al.
Published: (2025)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
by: Newman, Benjamin A., et al.
Published: (2024)
by: Newman, Benjamin A., et al.
Published: (2024)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
by: Ding, Hongyu, et al.
Published: (2026)
by: Ding, Hongyu, et al.
Published: (2026)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
by: Fang, Yu, et al.
Published: (2025)
by: Fang, Yu, et al.
Published: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
PokeFlex: A Real-World Dataset of Volumetric Deformable Objects for Robotics
by: Obrist, Jan, et al.
Published: (2024)
by: Obrist, Jan, et al.
Published: (2024)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
by: Zhang, Hai, et al.
Published: (2026)
by: Zhang, Hai, et al.
Published: (2026)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
by: Englmeier, Stefan, et al.
Published: (2026)
by: Englmeier, Stefan, et al.
Published: (2026)
Self-Supervised Online Robot-Agnostic Traversability Estimation for Open-World Environments
by: Hindel, Julia, et al.
Published: (2026)
by: Hindel, Julia, et al.
Published: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
by: Wen, Junjie, et al.
Published: (2024)
by: Wen, Junjie, et al.
Published: (2024)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
by: Wu, Yuhan, et al.
Published: (2025)
by: Wu, Yuhan, et al.
Published: (2025)
Similar Items
-
Playful DoggyBot: Learning Agile and Precise Quadrupedal Locomotion
by: Duan, Xin, et al.
Published: (2024) -
FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
by: Liu, Weiheng, et al.
Published: (2025) -
HumanPlus: Humanoid Shadowing and Imitation from Humans
by: Fu, Zipeng, et al.
Published: (2024) -
Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
by: Fu, Zipeng, et al.
Published: (2024) -
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
by: Cheng, An-Chieh, et al.
Published: (2024)