Do Visual Imaginations Improve Vision-and-Language Navigation Agents?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Perincherry, Akhil, Krantz, Jacob, Lee, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913748263895040
author Perincherry, Akhil
Krantz, Jacob
Lee, Stefan
author_facet Perincherry, Akhil
Krantz, Jacob
Lee, Stefan
contents Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations of sub-goals implied by the instructions can serve as navigational cues and lead to increased navigation performance. To synthesize these visual representations or imaginations, we leverage a text-to-image diffusion model on landmark references contained in segmented instructions. These imaginations are provided to VLN agents as an added modality to act as landmark cues and an auxiliary loss is added to explicitly encourage relating these with their corresponding referring expressions. Our findings reveal an increase in success rate (SR) of around 1 point and up to 0.5 points in success scaled by inverse path length (SPL) across agents. These results suggest that the proposed approach reinforces visual understanding compared to relying on language instructions alone. Code and data for our work can be found at https://www.akhilperincherry.com/VLN-Imagine-website/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16394
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
Perincherry, Akhil
Krantz, Jacob
Lee, Stefan
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Robotics
Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations of sub-goals implied by the instructions can serve as navigational cues and lead to increased navigation performance. To synthesize these visual representations or imaginations, we leverage a text-to-image diffusion model on landmark references contained in segmented instructions. These imaginations are provided to VLN agents as an added modality to act as landmark cues and an auxiliary loss is added to explicitly encourage relating these with their corresponding referring expressions. Our findings reveal an increase in success rate (SR) of around 1 point and up to 0.5 points in success scaled by inverse path length (SPL) across agents. These results suggest that the proposed approach reinforces visual understanding compared to relying on language instructions alone. Code and data for our work can be found at https://www.akhilperincherry.com/VLN-Imagine-website/.
title Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Robotics
url https://arxiv.org/abs/2503.16394