Continual Vision-and-Language Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeong, Seongjun, Kang, Gi-Cheon, Choi, Seongho, Kim, Joochan, Zhang, Byoung-Tak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914125317144576
author Jeong, Seongjun
Kang, Gi-Cheon
Choi, Seongho
Kim, Joochan
Zhang, Byoung-Tak
author_facet Jeong, Seongjun
Kang, Gi-Cheon
Choi, Seongho
Kim, Joochan
Zhang, Byoung-Tak
contents Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter novel environments. To address this, we propose the Continual Vision-and-Language Navigation (CVLN) paradigm, where agents learn and adapt incrementally across multiple \textit{scene domains}. CVLN includes two setups: Initial-instruction based CVLN for instruction-following, and Dialogue-based CVLN for dialogue-guided navigation. We also introduce two simple yet effective baselines for sequential decision-making: Perplexity Replay (PerpR), which replays difficult episodes, and Episodic Self-Replay (ESR), which stores and revisits action logits during training. Experiments show that existing continual learning methods fall short for CVLN, while PerpR and ESR achieve better performance by efficiently utilizing replay memory.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continual Vision-and-Language Navigation
Jeong, Seongjun
Kang, Gi-Cheon
Choi, Seongho
Kim, Joochan
Zhang, Byoung-Tak
Computer Vision and Pattern Recognition
Artificial Intelligence
Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter novel environments. To address this, we propose the Continual Vision-and-Language Navigation (CVLN) paradigm, where agents learn and adapt incrementally across multiple \textit{scene domains}. CVLN includes two setups: Initial-instruction based CVLN for instruction-following, and Dialogue-based CVLN for dialogue-guided navigation. We also introduce two simple yet effective baselines for sequential decision-making: Perplexity Replay (PerpR), which replays difficult episodes, and Episodic Self-Replay (ESR), which stores and revisits action logits during training. Experiments show that existing continual learning methods fall short for CVLN, while PerpR and ESR achieve better performance by efficiently utilizing replay memory.
title Continual Vision-and-Language Navigation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.15049