Harnessing Input-Adaptive Inference for Efficient VLN

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Dongwoo, Perincherry, Akhil, Coalson, Zachary, Gabriel, Aiden, Lee, Stefan, Hong, Sanghyun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912535657054208
author Kang, Dongwoo
Perincherry, Akhil
Coalson, Zachary
Gabriel, Aiden
Lee, Stefan
Hong, Sanghyun
author_facet Kang, Dongwoo
Perincherry, Akhil
Coalson, Zachary
Gabriel, Aiden
Lee, Stefan
Hong, Sanghyun
contents An emerging paradigm in vision-and-language navigation (VLN) is the use of history-aware multi-modal transformer models. Given a language instruction, these models process observation and navigation history to predict the most appropriate action for an agent. While they have significantly improved performance, the scale of these models can be a bottleneck in practical settings with limited computational resources. In this work, we propose a novel input-adaptive navigation method to enhance VLN model efficiency. We first show that existing input-adaptive mechanisms fail to reduce computations without substantial performance degradation. To address this, we introduce three adaptive algorithms, each deployed at a different level: (1) To improve spatial efficiency, we selectively process panoramic views at each observation of an agent. (2) To improve intra-model efficiency, we propose importance-based adaptive thresholding for the early-exit methods. (3) To improve temporal efficiency, we implement a caching mechanism that prevents reprocessing of views previously seen by the agent. In evaluations on seven VLN benchmarks, we demonstrate over a 2$\times$ reduction in computation across three off-the-shelf agents in both standard and continuous environments. Our code is publicly available at https://github.com/secure-ai-systems-group/adaptive-vision-and-language-navigation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Harnessing Input-Adaptive Inference for Efficient VLN
Kang, Dongwoo
Perincherry, Akhil
Coalson, Zachary
Gabriel, Aiden
Lee, Stefan
Hong, Sanghyun
Computer Vision and Pattern Recognition
Machine Learning
An emerging paradigm in vision-and-language navigation (VLN) is the use of history-aware multi-modal transformer models. Given a language instruction, these models process observation and navigation history to predict the most appropriate action for an agent. While they have significantly improved performance, the scale of these models can be a bottleneck in practical settings with limited computational resources. In this work, we propose a novel input-adaptive navigation method to enhance VLN model efficiency. We first show that existing input-adaptive mechanisms fail to reduce computations without substantial performance degradation. To address this, we introduce three adaptive algorithms, each deployed at a different level: (1) To improve spatial efficiency, we selectively process panoramic views at each observation of an agent. (2) To improve intra-model efficiency, we propose importance-based adaptive thresholding for the early-exit methods. (3) To improve temporal efficiency, we implement a caching mechanism that prevents reprocessing of views previously seen by the agent. In evaluations on seven VLN benchmarks, we demonstrate over a 2$\times$ reduction in computation across three off-the-shelf agents in both standard and continuous environments. Our code is publicly available at https://github.com/secure-ai-systems-group/adaptive-vision-and-language-navigation.
title Harnessing Input-Adaptive Inference for Efficient VLN
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2508.09262