"Where am I?" Scene Retrieval with Language

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Jiaqi, Barath, Daniel, Armeni, Iro, Pollefeys, Marc, Blum, Hermann
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917830874628096
author Chen, Jiaqi
Barath, Daniel
Armeni, Iro
Pollefeys, Marc
Blum, Hermann
author_facet Chen, Jiaqi
Barath, Daniel
Armeni, Iro
Pollefeys, Marc
Blum, Hermann
contents Natural language interfaces to embodied AI are becoming more ubiquitous in our daily lives. This opens up further opportunities for language-based interaction with embodied agents, such as a user verbally instructing an agent to execute some task in a specific location. For example, "put the bowls back in the cupboard next to the fridge" or "meet me at the intersection under the red sign." As such, we need methods that interface between natural language and map representations of the environment. To this end, we explore the question of whether we can use an open-set natural language query to identify a scene represented by a 3D scene graph. We define this task as "language-based scene-retrieval" and it is closely related to "coarse-localization," but we are instead searching for a match from a collection of disjoint scenes and not necessarily a large-scale continuous map. We present Text2SceneGraphMatcher, a "scene-retrieval" pipeline that learns joint embeddings between text descriptions and scene graphs to determine if they are a match. The code, trained models, and datasets will be made public.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14565
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle "Where am I?" Scene Retrieval with Language
Chen, Jiaqi
Barath, Daniel
Armeni, Iro
Pollefeys, Marc
Blum, Hermann
Computer Vision and Pattern Recognition
Natural language interfaces to embodied AI are becoming more ubiquitous in our daily lives. This opens up further opportunities for language-based interaction with embodied agents, such as a user verbally instructing an agent to execute some task in a specific location. For example, "put the bowls back in the cupboard next to the fridge" or "meet me at the intersection under the red sign." As such, we need methods that interface between natural language and map representations of the environment. To this end, we explore the question of whether we can use an open-set natural language query to identify a scene represented by a 3D scene graph. We define this task as "language-based scene-retrieval" and it is closely related to "coarse-localization," but we are instead searching for a match from a collection of disjoint scenes and not necessarily a large-scale continuous map. We present Text2SceneGraphMatcher, a "scene-retrieval" pipeline that learns joint embeddings between text descriptions and scene graphs to determine if they are a match. The code, trained models, and datasets will be made public.
title "Where am I?" Scene Retrieval with Language
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.14565