Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zehao, Wu, Minye, Cao, Yixin, Ma, Yubo, Chen, Meiqi, Tuytelaars, Tinne
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910620371124224
author Wang, Zehao
Wu, Minye
Cao, Yixin
Ma, Yubo
Chen, Meiqi
Tuytelaars, Tinne
author_facet Wang, Zehao
Wu, Minye
Cao, Yixin
Ma, Yubo
Chen, Meiqi
Tuytelaars, Tinne
contents This study presents a novel evaluation framework for the Vision-Language Navigation (VLN) task. It aims to diagnose current models for various instruction categories at a finer-grained level. The framework is structured around the context-free grammar (CFG) of the task. The CFG serves as the basis for the problem decomposition and the core premise of the instruction categories design. We propose a semi-automatic method for CFG construction with the help of Large-Language Models (LLMs). Then, we induct and generate data spanning five principal instruction categories (i.e. direction change, landmark recognition, region recognition, vertical movement, and numerical comprehension). Our analysis of different models reveals notable performance discrepancies and recurrent issues. The stagnation of numerical comprehension, heavy selective biases over directional concepts, and other interesting findings contribute to the development of future language-guided navigation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
Wang, Zehao
Wu, Minye
Cao, Yixin
Ma, Yubo
Chen, Meiqi
Tuytelaars, Tinne
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
This study presents a novel evaluation framework for the Vision-Language Navigation (VLN) task. It aims to diagnose current models for various instruction categories at a finer-grained level. The framework is structured around the context-free grammar (CFG) of the task. The CFG serves as the basis for the problem decomposition and the core premise of the instruction categories design. We propose a semi-automatic method for CFG construction with the help of Large-Language Models (LLMs). Then, we induct and generate data spanning five principal instruction categories (i.e. direction change, landmark recognition, region recognition, vertical movement, and numerical comprehension). Our analysis of different models reveals notable performance discrepancies and recurrent issues. The stagnation of numerical comprehension, heavy selective biases over directional concepts, and other interesting findings contribute to the development of future language-guided navigation systems.
title Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2409.17313