LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Tianrui, Yang, Yurou, Cheng, Harry, Lin, Muyuan, Kim, Richard, Madhivanan, Rajasimman, Sen, Arnie, Manocha, Dinesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929337948700672
author Guan, Tianrui
Yang, Yurou
Cheng, Harry
Lin, Muyuan
Kim, Richard
Madhivanan, Rajasimman
Sen, Arnie
Manocha, Dinesh
author_facet Guan, Tianrui
Yang, Yurou
Cheng, Harry
Lin, Muyuan
Kim, Richard
Madhivanan, Rajasimman
Sen, Arnie
Manocha, Dinesh
contents In this paper, we present LOC-ZSON, a novel Language-driven Object-Centric image representation for object navigation task within complex scenes. We propose an object-centric image representation and corresponding losses for visual-language model (VLM) fine-tuning, which can handle complex object-level queries. In addition, we design a novel LLM-based augmentation and prompt templates for stability during training and zero-shot inference. We implement our method on Astro robot and deploy it in both simulated and real-world environments for zero-shot object navigation. We show that our proposed method can achieve an improvement of 1.38 - 13.38% in terms of text-to-image recall on different benchmark settings for the retrieval task. For object navigation, we show the benefit of our approach in simulation and real world, showing 5% and 16.67% improvement in terms of navigation success rate, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
Guan, Tianrui
Yang, Yurou
Cheng, Harry
Lin, Muyuan
Kim, Richard
Madhivanan, Rajasimman
Sen, Arnie
Manocha, Dinesh
Computer Vision and Pattern Recognition
Robotics
In this paper, we present LOC-ZSON, a novel Language-driven Object-Centric image representation for object navigation task within complex scenes. We propose an object-centric image representation and corresponding losses for visual-language model (VLM) fine-tuning, which can handle complex object-level queries. In addition, we design a novel LLM-based augmentation and prompt templates for stability during training and zero-shot inference. We implement our method on Astro robot and deploy it in both simulated and real-world environments for zero-shot object navigation. We show that our proposed method can achieve an improvement of 1.38 - 13.38% in terms of text-to-image recall on different benchmark settings for the retrieval task. For object navigation, we show the benefit of our approach in simulation and real world, showing 5% and 16.67% improvement in terms of navigation success rate, respectively.
title LOC-ZSON: Language-driven Object-Centric Zero-Shot Object Retrieval and Navigation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2405.05363