MUVLA: Learning to Explore Object Navigation via Map Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Peilong, Jia, Fan, Zhang, Min, Qiu, Yutao, Tang, Hongyao, Zheng, Yan, Wang, Tiancai, Hao, Jianye
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908568588910592
author Han, Peilong
Jia, Fan
Zhang, Min
Qiu, Yutao
Tang, Hongyao
Zheng, Yan
Wang, Tiancai
Hao, Jianye
author_facet Han, Peilong
Jia, Fan
Zhang, Min
Qiu, Yutao
Tang, Hongyao
Zheng, Yan
Wang, Tiancai
Hao, Jianye
contents In this paper, we present MUVLA, a Map Understanding Vision-Language-Action model tailored for object navigation. It leverages semantic map abstractions to unify and structure historical information, encoding spatial context in a compact and consistent form. MUVLA takes the current and history observations, as well as the semantic map, as inputs and predicts the action sequence based on the description of goal object. Furthermore, it amplifies supervision through reward-guided return modeling based on dense short-horizon progress signals, enabling the model to develop a detailed understanding of action value for reward maximization. MUVLA employs a three-stage training pipeline: learning map-level spatial understanding, imitating behaviors from mixed-quality demonstrations, and reward amplification. This strategy allows MUVLA to unify diverse demonstrations into a robust spatial representation and generate more rational exploration strategies. Experiments on HM3D and Gibson benchmarks demonstrate that MUVLA achieves great generalization and learns effective exploration behaviors even from low-quality or partially successful trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25966
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MUVLA: Learning to Explore Object Navigation via Map Understanding
Han, Peilong
Jia, Fan
Zhang, Min
Qiu, Yutao
Tang, Hongyao
Zheng, Yan
Wang, Tiancai
Hao, Jianye
Robotics
In this paper, we present MUVLA, a Map Understanding Vision-Language-Action model tailored for object navigation. It leverages semantic map abstractions to unify and structure historical information, encoding spatial context in a compact and consistent form. MUVLA takes the current and history observations, as well as the semantic map, as inputs and predicts the action sequence based on the description of goal object. Furthermore, it amplifies supervision through reward-guided return modeling based on dense short-horizon progress signals, enabling the model to develop a detailed understanding of action value for reward maximization. MUVLA employs a three-stage training pipeline: learning map-level spatial understanding, imitating behaviors from mixed-quality demonstrations, and reward amplification. This strategy allows MUVLA to unify diverse demonstrations into a robust spatial representation and generate more rational exploration strategies. Experiments on HM3D and Gibson benchmarks demonstrate that MUVLA achieves great generalization and learns effective exploration behaviors even from low-quality or partially successful trajectories.
title MUVLA: Learning to Explore Object Navigation via Map Understanding
topic Robotics
url https://arxiv.org/abs/2509.25966