NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wanchoo, Karan, Zuo, Xiaoye, Gonzalez, Hannah, Dan, Soham, Georgakis, Georgios, Roth, Dan, Daniilidis, Kostas, Miltsakaki, Eleni
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910750562320384
author Wanchoo, Karan
Zuo, Xiaoye
Gonzalez, Hannah
Dan, Soham
Georgakis, Georgios
Roth, Dan
Daniilidis, Kostas
Miltsakaki, Eleni
author_facet Wanchoo, Karan
Zuo, Xiaoye
Gonzalez, Hannah
Dan, Soham
Georgakis, Georgios
Roth, Dan
Daniilidis, Kostas
Miltsakaki, Eleni
contents We present NAVCON, a large-scale annotated Vision-Language Navigation (VLN) corpus built on top of two popular datasets (R2R and RxR). The paper introduces four core, cognitively motivated and linguistically grounded, navigation concepts and an algorithm for generating large-scale silver annotations of naturally occurring linguistic realizations of these concepts in navigation instructions. We pair the annotated instructions with video clips of an agent acting on these instructions. NAVCON contains 236, 316 concept annotations for approximately 30, 0000 instructions and 2.7 million aligned images (from approximately 19, 000 instructions) showing what the agent sees when executing an instruction. To our knowledge, this is the first comprehensive resource of navigation concepts. We evaluated the quality of the silver annotations by conducting human evaluation studies on NAVCON samples. As further validation of the quality and usefulness of the resource, we trained a model for detecting navigation concepts and their linguistic realizations in unseen instructions. Additionally, we show that few-shot learning with GPT-4o performs well on this task using large-scale silver annotations of NAVCON.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13026
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
Wanchoo, Karan
Zuo, Xiaoye
Gonzalez, Hannah
Dan, Soham
Georgakis, Georgios
Roth, Dan
Daniilidis, Kostas
Miltsakaki, Eleni
Computation and Language
Computer Vision and Pattern Recognition
We present NAVCON, a large-scale annotated Vision-Language Navigation (VLN) corpus built on top of two popular datasets (R2R and RxR). The paper introduces four core, cognitively motivated and linguistically grounded, navigation concepts and an algorithm for generating large-scale silver annotations of naturally occurring linguistic realizations of these concepts in navigation instructions. We pair the annotated instructions with video clips of an agent acting on these instructions. NAVCON contains 236, 316 concept annotations for approximately 30, 0000 instructions and 2.7 million aligned images (from approximately 19, 000 instructions) showing what the agent sees when executing an instruction. To our knowledge, this is the first comprehensive resource of navigation concepts. We evaluated the quality of the silver annotations by conducting human evaluation studies on NAVCON samples. As further validation of the quality and usefulness of the resource, we trained a model for detecting navigation concepts and their linguistic realizations in unseen instructions. Additionally, we show that few-shot learning with GPT-4o performs well on this task using large-scale silver annotations of NAVCON.
title NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13026