HomeRobot: Open-Vocabulary Mobile Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Yenamandra, Sriram, Ramachandran, Arun, Yadav, Karmesh, Wang, Austin, Khanna, Mukul, Gervet, Theophile, Yang, Tsung-Yen, Jain, Vidhi, Clegg, Alexander William, Turner, John, Kira, Zsolt, Savva, Manolis, Chang, Angel, Chaplot, Devendra Singh, Batra, Dhruv, Mottaghi, Roozbeh, Bisk, Yonatan, Paxton, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
by: Khanna, Mukul, et al.
Published: (2024)
by: Khanna, Mukul, et al.
Published: (2024)
Towards Open-World Mobile Manipulation in Homes: Lessons from the Neurips 2023 HomeRobot Open Vocabulary Mobile Manipulation Challenge
by: Yenamandra, Sriram, et al.
Published: (2024)
by: Yenamandra, Sriram, et al.
Published: (2024)
Situated Instruction Following
by: Min, So Yeon, et al.
Published: (2024)
by: Min, So Yeon, et al.
Published: (2024)
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
by: Jain, Vidhi, et al.
Published: (2024)
by: Jain, Vidhi, et al.
Published: (2024)
ReLIC: A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI
by: Elawady, Ahmad, et al.
Published: (2024)
by: Elawady, Ahmad, et al.
Published: (2024)
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control
by: Gupta, Gunshi, et al.
Published: (2024)
by: Gupta, Gunshi, et al.
Published: (2024)
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning
by: Ramrakhya, Ram, et al.
Published: (2025)
by: Ramrakhya, Ram, et al.
Published: (2025)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
by: Yadav, Karmesh, et al.
Published: (2025)
by: Yadav, Karmesh, et al.
Published: (2025)
Controllable Human-Object Interaction Synthesis
by: Li, Jiaman, et al.
Published: (2023)
by: Li, Jiaman, et al.
Published: (2023)
VisACD: Visibility-Based GPU-Accelerated Approximate Convex Decomposition
by: Fokin, Egor, et al.
Published: (2026)
by: Fokin, Egor, et al.
Published: (2026)
Learning Model Successors
by: Chang, Yingshan, et al.
Published: (2025)
by: Chang, Yingshan, et al.
Published: (2025)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Seeing the Unseen: Visual Common Sense for Semantic Placement
by: Ramrakhya, Ram, et al.
Published: (2024)
by: Ramrakhya, Ram, et al.
Published: (2024)
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
by: Andrade, Moises, et al.
Published: (2025)
by: Andrade, Moises, et al.
Published: (2025)
DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement
by: Newman, Benjamin A., et al.
Published: (2024)
by: Newman, Benjamin A., et al.
Published: (2024)
EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates
by: Peng, Weikun, et al.
Published: (2026)
by: Peng, Weikun, et al.
Published: (2026)
Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
by: Soraki, Rustin, et al.
Published: (2026)
by: Soraki, Rustin, et al.
Published: (2026)
Gradient Localization Improves Lifelong Pretraining of Language Models
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Survey on Modeling of Human‐made Articulated Objects
by: Jiayi Liu, et al.
Published: (2025)
by: Jiayi Liu, et al.
Published: (2025)
Survey on Modeling of Human-made Articulated Objects
by: Liu, Jiayi, et al.
Published: (2024)
by: Liu, Jiayi, et al.
Published: (2024)
REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
by: Thompson, Jacob, et al.
Published: (2025)
by: Thompson, Jacob, et al.
Published: (2025)
SOC Increase in UK Topsoils Is Most Likely due to SOC Vertical Redistribution: Comment on Bentley et al. (2025)
by: Vincent Chaplot
Published: (2026)
by: Vincent Chaplot
Published: (2026)
Text-to-3D Shape Generation
by: Lee, Han-Hung, et al.
Published: (2024)
by: Lee, Han-Hung, et al.
Published: (2024)
iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos
by: Peng, Weikun, et al.
Published: (2025)
by: Peng, Weikun, et al.
Published: (2025)
WildSmoke: Ready-to-Use Dynamic 3D Smoke Assets from a Single Video in the Wild
by: Liu, Yuqiu, et al.
Published: (2025)
by: Liu, Yuqiu, et al.
Published: (2025)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
Generalizing Single-View 3D Shape Retrieval to Occlusions and Unseen Objects
by: Wu, Qirui, et al.
Published: (2023)
by: Wu, Qirui, et al.
Published: (2023)
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
by: Hu, Yafei, et al.
Published: (2023)
by: Hu, Yafei, et al.
Published: (2023)
Environmental geopolitics of the Caspian basin energy interactions
by: Mottaghi, A.
Published: (2016)
by: Mottaghi, A.
Published: (2016)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024)
by: Sun, Jimin, et al.
Published: (2024)
CAGE: Controllable Articulation GEneration
by: Liu, Jiayi, et al.
Published: (2023)
by: Liu, Jiayi, et al.
Published: (2023)
Cover crop studies: The need for more reliable data
by: Vincent Chaplot, et al.
Published: (2024)
by: Vincent Chaplot, et al.
Published: (2024)
Diorama: Unleashing Zero-shot Single-view 3D Indoor Scene Modeling
by: Wu, Qirui, et al.
Published: (2024)
by: Wu, Qirui, et al.
Published: (2024)
R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
by: Wu, Qirui, et al.
Published: (2024)
by: Wu, Qirui, et al.
Published: (2024)
Similar Items
-
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
by: Khanna, Mukul, et al.
Published: (2024) -
Towards Open-World Mobile Manipulation in Homes: Lessons from the Neurips 2023 HomeRobot Open Vocabulary Mobile Manipulation Challenge
by: Yenamandra, Sriram, et al.
Published: (2024) -
Situated Instruction Following
by: Min, So Yeon, et al.
Published: (2024) -
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
by: Jain, Vidhi, et al.
Published: (2024) -
ReLIC: A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI
by: Elawady, Ahmad, et al.
Published: (2024)